Hepatocellular carcinoma (HCC) is the most common primary liver cancer and represents one of the leading causes of cancer‑related mortality worldwide [1]. Advances in systemic therapy over the last decade have transformed the therapeutic landscape of advanced HCC [2–4]. Historically, treatment options were limited to multikinase inhibitors [5–7]. The emergence of immune checkpoint inhibitors (ICI) and combination strategies has substantially improved outcomes for patients with advanced HCC. Several pivotal randomized clinical trials (RCT) have redefined the first‑line treatment paradigm including IMbrave‑150[8,9]HIMALAYA [10], LEAP‑002 [11], CARES‑310 [12], and CheckMate‑9DW [13], among others [14] (Table 1). Other RCTs attempted therapeutic combinations, either with cabozantinib [15] or lenvatinib [11], not showing clinical benefit. More recently, although promising phase II trials [16], the combination with anti-PD-L1 and anti-TIGIT (tiragolumab) failed to demonstrate additional clinical benefit (IMbrave-152 NCT05904886). Other trials have been conducted mostly in Asian countries needing validation in Western populations [17,18]Moreover, combined immunotherapy with or without locoregional therapies, showed prolonged progression-free survival (PFS), with overall survival (OS) benefit still pending [19,20].
Most relevant worldwide phase III, first-line, systemic therapy trials for advanced HCC. Population and demographic characteristics.
Note: Data shown from trials publications. IMbrave-150 [8,9], HIMALAYA [10], LEAP-002 [11], CARES-310 [12], COSMIC [15], and CheckMate-9DW [13].
Abbreviations: AFP: alpha-fetoprotein; BCLC: Barcelona Clinic Liver Cancer; ECOG: performance status; EH: metastatic extrahepatic disease; HCC: hepatocellular carcinoma; HBV: chronic hepatitis B virus, HCV: chronic hepatitis C virus, MVI: macrovascular tumor vein invasion; PVT: portal vein tumor thrombosis; RCT: randomized clinical trial.
Early signals of efficacy have also been evaluated in the neo/adjuvant setting before ablation, surgical resection [21,22] or liver transplantation [23–26], but are outside the scope of this review. Early phase clinical trials have reported encouraging rates of major pathological response [27]. However, optimal patient selection, treatment duration, and safety considerations remain uncertain areas [28]. Meanwhile, in the adjuvant setting, two recent RCTs, the IMbrave-050 [29] and KEYNOTE-937 (NCT03867084), failed to show a reduction in the incidence of recurrent HCC after surgery or ablation therapies.
Despite these advances, interpreting results from these trials requires a strong methodological understanding of survival analysis [30], particularly for treatment effects under time‑to‑event outcomes [31]. However, interpreting these outcomes can be complex, particularly in HCC where underlying liver disease introduces competing risks and heterogeneous clinical trajectories [14]. As a result, clinicians must interpret trial results not only in terms of statistical significance but also regarding methodological validity and clinical relevance.
HCC arises in the context of chronic liver disease in approximately 90–95% of patients, with cirrhosis in 70–90% of the cases. Consequently, prognosis is not determined solely by tumor burden but also by the severity of the underlying liver dysfunction. Decompensation of cirrhosis significantly worsens outcomes and may limit treatment eligibility, tolerance, and therapeutic response [31,32]. In clinical practice, inappropriate treatment selection or persistence with ineffective therapies may precipitate hepatic decompensation, thereby negatively affecting survival independently of tumor control (Fig. 1). These factors illustrate the complexity of interpreting therapeutic outcomes in HCC.
2BCLC staging for clinical decision-makingThe Barcelona Clinic Liver Cancer (BCLC) staging system represents the most widely adopted framework for integrating tumor burden, liver function, and patient performance status to guide treatment decisions in hepatocellular carcinoma [33]. The latest update recognizes substantial heterogeneity within stages emphasizing a more individualized approach to treatment selection under uncertain scenarios of evidence (CUSE framework) [34]. The updated BCLC recommendations also emphasize concepts such as “treatment stage migration” [35,36] and “treatable tumor regression” [37–39]. These principles acknowledge that patients may benefit from therapies typically recommended for different stages when individual clinical circumstances justify alternative approaches. Such flexibility is particularly relevant in the era of evolving systemic therapies and combined treatment strategies.
3Correlation, association and causal associationsFirst, it is essential to distinguish correlation, association, and causation. Although all these terms seem to be rather similar, they are not. Correlation refers to the degree to which two continuous variables vary together (variability of y in terms of x, and vice versa), usually without implying any directional or mechanistic relationship. Spearman or Pearson’s statistical tests are often conducted to test whether there is a positive or negative correlation, meaning that the variability of y is somehow explained by a linear relation with the variability of x. These statistical tests have a high power to test minimal deviations of the null hypothesis. Consequently, P values do not account for the magnitude of the effect, rather the “R” or “R2” provide a measure of a magnitude of this linear relationship between variables (variability). Association is a broader concept indicating that two variables are statistically related, but the relationship may reflect confounding, selection bias, or reverse causality. Causation, in contrast, implies that changing one variable produces a change in the outcome. In other words, correlation or association does not mean causation.
RCTs are specifically designed to support causal inference because randomization balances measured and unmeasured confounders between treatment groups [40,41]. Therefore, the difference in overall survival between treatment arms may be interpreted as causal within the study population, only if there were no bias in the execution of the trial. By contrast, subgroup findings, post hoc analyses, and indirect comparisons across different trials should usually be interpreted as associative rather than causal [42].
4Hierarchy of endpoints in HCC trialsEndpoints represent measurable outcomes used to evaluate treatment effects. In oncology trials, endpoints can be categorized as clinical endpoints or surrogate endpoints [43,44]. Clinical endpoints directly reflect patient benefit, whereas surrogate endpoints are intermediate outcomes. Overall survival (OS) remains the gold standard endpoint because it directly measures the goal of therapy: prolongation of life. However, OS may require prolonged follow‑up and large sample sizes. Consequently, surrogate endpoints such as progression‑free survival (PFS), time‑to‑progression (TTP), and objective response rate (ORR) are frequently used to accelerate drug development or regulatory approval. In HCC, the selection of endpoints is particularly complex due to the interplay between tumor progression and liver dysfunction. Death may occur due to tumor progression, complications of cirrhosis, or unrelated causes. This complexity underscores the need for careful interpretation of surrogate outcomes [45].
A surrogate endpoint is considered valid if it reliably predicts the effect of treatment on a true clinical endpoint. Validation requires evidence of correlation at both the individual and trial levels [46,47]. Several studies have examined the relationship between PFS and OS in oncology trials, showing moderate, weak or inconsistent correlations [46,47]. In HCC, surrogate endpoint validation is complicated by tumor heterogeneity across patients and etiologies, or subsequent therapies after disease progression [44].
The design of clinical trials includes phase I-II and III trials. In phase II trials, the main objective is not only safety (particularly for phase I studies), but also early signs of clinical efficacy including ORR or PFS. Moreover, in early or intermediate stage HCC trials, the primary objective could be recurrence-free survival or PFS, respectively. On the other hand, phase III clinical trials in advanced HCC attempt to evaluate OS, either as superiority or non-inferiority designs. This distinction should be taken into consideration to adequately interpret outcome measures, sample sizes, and survival curves. The expected “non-inferior” outcome depends on the expected limit on worse outcome that the investigators allow the intervention compared to the standard of care. As the expected clinical differences are rather tight or small, sample sizes tend to be much larger than superiority designs.
5Prognostic variables in HCC trials: the central role of liver function in systemic therapyPrognostic variables are factors associated with patient outcomes independently of treatment, whereas predictive variables modify the effect of a specific therapy and help identify patients most likely to benefit from an intervention. In HCC, prognosis is determined by a complex interaction between tumor characteristics, liver function, and patient-related factors [33]. Tumor burden—including number of lesions, maximum tumor diameter, vascular invasion, and extrahepatic spread—represents a major determinant of survival. Tumor biology also plays a critical role, with biomarkers such as alpha-fetoprotein (AFP) providing additional prognostic information [48–50]. Patient-related factors including age, comorbidities, and performance status (ECOG) further contribute to the overall prognosis. Recognizing these variables is essential when interpreting randomized trials, as differences in baseline patient characteristics may influence outcomes and complicate comparisons across studies.
Liver function represents a critical determinant of both treatment eligibility and survival outcomes [51]. Variables such as Child–Pugh class, MELD score, ALBI grade, and the presence of clinically significant portal hypertension influence both survival and treatment eligibility [33]. More granular assessments of liver function, such as the albumin–bilirubin (ALBI) score, may provide additional prognostic information within Child–Pugh A patients [2,52].
However, liver function is not static and therefore should be considered a dynamic process. Hepatic decompensation may occur during systemic therapy and is associated with significantly worse outcomes [36,53–56]. Recent studies have demonstrated that early hepatic decompensation is a strong predictor of mortality in patients receiving systemic therapy, independently from tumor control [53,56,57]. More importantly, cirrhosis recompensation may be achieved after liver injury eradication [58,59], further demanding a much more complex clinical decision-making process (Fig. 2).
6Interpreting survival analysis in time-to-event outcomes: proportional and non-proportional hazardsTime‑to‑event outcomes are typically analyzed using Kaplan–Meier estimator of survival, a non‑parametric estimator used to calculate survival probabilities over time. The estimator accounts for censored observations, individuals who have not yet experienced the event of interest at the time of analysis or who were lost to follow‑up. The Kaplan–Meier survival function [S(t)] is estimated as the product of conditional survival probabilities across successive time intervals and can be graphically plotted on survival curves (Fig. 3). Each step in the survival curve represents an event occurrence, while horizontal segments represent periods without events. This method allows visualization of the dynamic occurrence of events during follow‑up.
A Kaplan-Meier survival curve.
On the contrary, the hazard function represents the instantaneous risk of experiencing the event among individuals who have not yet experienced it [60,61]. Unlike probability, the hazard describes how likely events occur over time (instantaneous risk), and the hazard ratios (HR) derived from the Cox regression model provide only a summary of treatment effect and rely on Cox model assumptions. The effect measure or HR (the ratio between hazards of groups) is that instantaneous risk on average or as a summary over the entire follow-up period. However, this “average” effect should be interpreted carefully, and not independently of Kaplan-Meier curves (Fig. 3). As a relative effect outcome measure, a HR below one indicates reduced risk with treatment (or exposed group), whereas above one indicates increased risk.
The Cox proportional hazards regression model is the most commonly used mathematical model for analyzing time‑to‑event data [62–64]. The model estimates HR while allowing adjustment for covariates. A key assumption of the model is “proportional hazards”, meaning that the HR between treatment groups remains “constant” throughout the entire follow‑up period [61,65,66].
The proportional hazards assumption can be evaluated using several methods, statistically and graphically. Schoenfeld residuals are commonly used to assess whether hazard ratios change over time. If residuals display systematic patterns with respect to time, the assumption may be violated. Graphical inspection of survival curves also provides valuable insight. Parallel curves generally support the proportional hazards assumption, whereas crossing or diverging curves may indicate violations (Fig. 4).
An example of evaluating graphically the proportional hazard assumption.
Immunotherapy trials frequently demonstrate early or delayed treatment effects. In these cases, survival curves may overlap early and converge or diverge later during follow‑up (e.g., IMbrave-050 [29], or HIMALAYA trial [10]). Crossing curves may also occur when early events (death) due to toxicity or other treatment-unrelated causes are followed by long‑term benefit among responders (e.g., CheckMate-9DW [13]). These patterns violate the proportional hazards assumption and complicate interpretation of a single HR over the entire follow-up period. Alternative approaches such as restricted mean survival time (an absolute measure of effect) or weighted log‑rank tests may provide complementary insights [67,68]. Nevertheless, readers should interpret cautiously the meaning of that outcome measure, and what is the null hypothesis of that statistical test.
The CheckMate-9DW trial is particularly informative because it illustrates several of the main methodological challenges in interpreting modern immunotherapy trials (Table 1). Reported efficacy outcomes included a HR of 0.79 for OS, and no beneficial effect upon PFS. While ORR was markedly higher with dual immunotherapy (36%vs 13%), complete responses were also more frequent (7%vs 2%), and median duration of response (DOR) was substantially longer (30.4 vs 12.9 months) [13]. This discordance between OS and PFS is not unusual with ICIs and was similarly observed in HIMALAYA [10] and LEAP-02 trials [11]. In other words, these regimens may not prevent early progression better than a tyrosine kinase inhibitor, but responses appear more meaningful and longer lasting, even when achieving stable disease. That is precisely why median PFS alone is inadequate to summarize the value of immunotherapy-based regimens in HCC. Moreover, not all types of progression are associated with worse OS [69].
CheckMate-9DW is a good example of why cross-trial comparisons are hazardous. Differences in control-arm choice, geographic distribution, HBV prevalence, rate of extrahepatic spread, ECOG distribution, AFP values, prevalent portal vein tumor invasion criteria (Vp1–4) or the proportion of BCLC-B patients can move median OS substantially [70]. Therefore, the median 23.7-month OS in CheckMate-9DW should not be interpreted as proof of superiority across trials. The trial also fits the broader discussion of non-proportional hazards in ICIs [71], in which a single HR can be clinically complex to interpret. If the survival curves separate late, the reported HR becomes a time-averaged summary that may dilute the magnitude of late benefit. This is especially relevant for dual checkpoint inhibition, where early toxicity, or treatment unrelated causes of death, or delayed immune activation may cause the treatment effect to vary over follow-up [72]. Thus, CheckMate-9DW is exactly the kind of trial where inspection of the full Kaplan–Meier curves, milestone survival, and tail-of-the-curve behavior would be clinically more informative than relying on median survival or an “on average” single HR estimate [66,67,73].
More importantly is to try to understand the causes of crossing Kaplan-Meier curves, and such deviation from the proportional hazard’s assumption. In other words, what was the cause of increased deaths over the first months with nivolumab plus ipilimumab? In this trial, an increased risk of death due to safety profile rather than tumor progression was observed, particularly liver-related causes of death [13]. Nevertheless, whether patients with cirrhosis and clinically significant portal hypertension (CSPH) are at an increased risk of death with ipilimumab plus nivolumab, remains uncertain. That interpretation is closer to real clinical practice than a simplistic reading of the HR alone.
7Interpretation of modern phase ii-iii clinical trials in HCC: interpreting median survival timesRecent RCT have significantly improved outcomes for advanced HCC (Table 1). However, comparisons across trials regarding median survival times should be interpreted cautiously. Understanding methodological differences between trials is essential for translating evidence into clinical practice.
The median overall survival is a time point in the Kaplan-Meier curve in which 50% of the study population or in each treatment arm (exposed or unexposed groups) still did not develop the event of interest (in this case, death). It gives information about the dynamics of the event of interest. However, it is not a robust outcome measure. As a continuous measure, median time is not “sensitive” to “outliers”, or long-term survivors. On the other hand, if the observed number of events is rather small, and 50% of the study population has not developed the event, median times cannot be estimated. Finally, comparing median survival times between groups is different from comparing other effect measures. In fact, 95% confidence intervals shown in each group would probably cross, and estimating P values using non-parametric statistic methods would not give sufficient power to address statistically significant differences in median survival times. Readers could review this key point observing that in most trials, 95% CI of median survival times cross-over between groups (Fig. 3).
Variations in the proportion of patients with macrovascular tumor invasion (Vp), extrahepatic spread, BCLC stage B, or high AFP levels, can significantly affect overall survival in the control arms of these studies. Moreover, the proportion of patients with cirrhosis, particularly those presenting CSPH, is another prognostic factor. Unfortunately, the proportion of patients with cirrhosis and CSPH was not clearly shown in most recent phase III clinical trials [74]. It is expected that patients with cirrhosis at baseline, particularly those with CSPH, may show increased risk of hepatic decompensation, and increased risk of liver-related mortality when compared to patients without cirrhosis [53].
8Translating trial evidence into clinical practiceA significant gap often exists between highly selected populations enrolled in RCT and those from routine clinical practice [75]. Most pivotal trials have enrolled patients with preserved liver function (Child–Pugh A), ECOG 0–1, and limited comorbidities. Also, we have previously detailed the exclusion criteria for patients with Vp4. In contrast, many patients in real-world settings present with more advanced liver dysfunction, portal hypertension, or multiple comorbid conditions. Clinical decision-making must therefore consider factors that are not always fully captured in trial populations, including the risk of bleeding, the presence of CSPH, cardiovascular comorbidities, and the potential for treatment-related hepatic decompensation.
Eligibility criteria should guide clinicians for external validity, or in other words, clinical applicability. In most of these first-line phase III trials, patients presenting Vp4 (HCC vascular tumor invasion of the main portal trunk) were excluded. The reason for excluding this group of patients in these trials is still uncertain and not convincing at all. In the IMbrave-150 study, 15% of total study population presented Vp4, but this baseline prognostic factor was not considered as a stratifying factor. However, treatment arms were balanced (Vp4 in atezo+bev 13.1% vs sorafenib 13.9%) and showed a trend towards clinical survival benefit within this particular sub-group of patients with a HR of 0.62 (not statistically significant probably due to small sample size and unprecise estimation) [76].
Defining cirrhosis and evaluating the risk of hepatic decompensation is essential at baseline. We have previously shown the proportion of cirrhosis, < 40% in the IMbrave-150 [74], and was recently reported in CheckMate‑9DW as 45.6% of the total study population (n = 305/668). In that trial, the definition of cirrhosis was retrospectively assessed using either histological METAVIR F4, or elastography > 14.6 kPa, or a FIB-4 > 3.25. However, a significant number of participants were indeterminate for these values [13]. So, the exact proportion of patients with cirrhosis is difficult to assess retrospectively in these trials.
Patients with HCC and CSPH are at increased risk of gastrointestinal bleeding with anti-angiogenic or anti-VEGF therapies. For this reason, it is important to assess the bleeding risk before treatment initiation. In the IMbrave-150 trial [77], patients were required to undergo upper endoscopy within six months prior to treatment initiation, and to start preventive strategies, either beta-blockers or endoscopic band ligation, for patients with high-risk varices. Nevertheless, in the clinical setting, most gastroenterologists or oncologists got confused with these eligibility criteria, changing prophylactic therapies for variceal bleeding, not fully detailed in the IMbrave-150 study protocol, and going against international consensus statements [51]. In fact, in that study, either beta-blockers or endoscopic band ligation could be done appropriately based on local or center-based decisions. These recommendations came in the pre-era of “beta-blockers” for all patients with cirrhosis (PREDESCI trial) [78,79], in which it has been shown that the use of carvedilol is associated with a reduction in the incidence of hepatic decompensation events. Consequently, it is not surprising that in the IMbrave-150 trial, a proportion of patients with small varices were untreated [77]. Despite these precautions, gastrointestinal bleeding remains a potential complication of anti-angiogenic or anti-VEGF therapies. However, reported rates of severe bleeding events in clinical trials have generally been low, highlighting the importance of careful patient selection [51]. Patients with Vp4 presenting large varices, with red spots at upper endoscopy, are at highest risk of variceal bleeding. Thus, this group of patients should be carefully evaluated, and the systemic treatment decision be appropriately conducted under multidisciplinary care settings.
9Systemic therapy trials and underrepresented patient populationsDespite the remarkable advances in systemic therapy for hepatocellular carcinoma, most RCT have enrolled highly selected patient populations. All the aforementioned trials excluded patients with Child–Pugh B cirrhosis, or with recent or prior hepatic decompensation (variceal bleeding, ascites, hepatic encephalopathy, bacterial peritonitis, among others). Nevertheless, it should be clarified that Child–Pugh A eligibility criteria has been extrapolated to non-cirrhotic patients, not adequately reported in trials. These restrictive criteria limit the generalizability of trial results to real-world populations, where patients often present with more complex clinical profiles.
As a result, clinicians must interpret trial results cautiously when applying them to broader patient populations. Real-world observational studies [54] and single-arm clinical trials [80,81] have therefore become increasingly important for evaluating treatment effectiveness and safety outside the strict conditions of RCT. Patients with cirrhosis showed significantly lower survival rates compared with non-cirrhotic patients [58]. Similarly, baseline Child–Pugh B liver function is associated with increased mortality risk following initiation of systemic therapy [81–85].
Although clinical trials generally do not impose strict age limits, elderly patients may present with additional comorbidities and altered pharmacokinetics that could influence treatment tolerability and increase the incidence of adverse events, although showing similar efficacy outcomes compared with younger patients. Patients with cardiovascular disease, renal impairment [86], autoimmune disorders, or prior organ transplantation present additional therapeutic challenges [87] ICIs may exacerbate autoimmune conditions or induce immune-mediated adverse events, while anti-angiogenic therapies may increase cardiovascular and bleeding risks. Patients with HIV infection or viral coinfections (HBV-HCV coinfection) are still underrepresented. Although available evidence suggests that ICIs do not significantly alter CD4+ counts or HIV viral load, additional studies are required to confirm long-term safety [88].
10Biomarkers and treatment responseBiomarkers are increasingly being investigated as tools to improve patient selection for systemic therapies in hepatocellular carcinoma. AFP remains the most widely used biomarker in clinical practice, with approximately 40% of patients with advanced HCC presenting with AFP levels greater than 400 ng/ml [89]. Changes in AFP levels during treatment may provide prognostic information. Early reductions in AFP have been associated with improved survival outcomes in several studies evaluating anti-angiogenic therapies and immunotherapy [49,90–92]. However, treatment response may occur independently of baseline AFP levels [92]. Additional biomarkers, including tumor mutational burden and circulating markers such as fibroblast growth factor 21 (FGF21) [93], are currently being investigated for their potential predictive role.
11The BCLC concept of “treatable tumor regression”: conversion therapyRecent advances in systemic therapy have introduced the concept of “conversion therapy” [39], in which initially unresectable hepatocellular carcinoma becomes amenable to potentially curative treatment after response to systemic therapy. This paradigm represents an important shift in the management of intermediate and advanced stage HCC. The BCLC update has integrated this concept into a more general novel term “treatable tumor regression” [33]. Reported conversion rates vary widely across studies, ranging from approximately 5% to 32% [39].
Importantly, however, curative conversion remains insufficiently captured in current clinical trial design in intermediate and advanced stage HCC. Despite increasing clinical recognition, most phase III trials do not prospectively define or systematically assess conversion to curative intent, and subsequent potentially curative local therapies are reported at relatively similar—and generally low—rates across studies [39]. This likely reflects, at least in part, a prevailing tendency to continue “effective” systemic therapy rather than actively pursuing alternative, potentially curative strategies once tumor regression is achieved. Moreover, there is currently no standardized definition of curative conversion, nor a globally accepted endpoint to capture this clinically meaningful outcome in trials, further limiting cross-study comparisons and the integration of this concept into evidence-based treatment algorithms.
These strategies highlight the importance of defining treatment upfront at the initiation of treatment and the continuous reassessment of tumor response during systemic therapy to identify patients who may become candidates for potentially curative treatment [23–26]. This dynamic treatment approach reflects the evolving nature of therapeutic decision-making in HCC, and probably new post-treatment prognostic allocation or staging systems may be further re-addressed soon (e.g., liver transplantation eligibility criteria) [36,94].
In some patients, continuation of immunotherapy beyond radiologic progression may be considered if such progression is not associated with a significant detriment in post-progression survival [36]. On the other hand, complete responders with ICIs may occasionally translate into long-term disease control or even “cure” in selected patients [37]. However, the definition of “cure” of HCC after immunotherapy is challenging, and ICIs treatment duration or withdrawal remains uncertain.
12Managing immune-related adverse events: etiologic factor, cirrhosis and portal hypertensionThe incidence and grades of immune-related adverse events in phase II and III clinical trials range according to the scheme (single versus dual immunotherapy), combined immunotherapy (PD-1, PD-L1 with or without anti-CTLA-4), and the total cycles of the induction phase (single dose anti-CTLA-4 vs more doses). Serious immune-related adverse events and high-dose steroids requirement range from 12% with atezolizumab + bevacizumab, 20% with durvalumab + tremelimumab, and almost 30% with nivolumab + ipilimumab. In the CheckMate-9DW trial [13], the dose of ipilimumab was relatively higher when compared to other solid tumors (3 mg/kg). This dose was selected in a prior phase II study (CheckMate-040), associated with better ORR and PFS, while preserving safety [95]. Although choosing a lower dose of ipilimumab has been tested in a phase II study with some signals of clinical efficacy (CHECKMATE-040), it is not approved, nor associated with clinical benefit in HCC in a phase III RCT. On the other hand, the HIMALAYA trial was designed exploring two different treatment arms, one including a single induction dose (tremelimumab 300 mg), and another with 4 sequential doses (tremelimumab 75 mg each) [10]. The latter arm was withdrawn due to not meaningful clinical efficacy in an interim analysis compared to durvalumab arm.
Most frequent immune-related adverse events are rash, pruritus, increasing liver enzymes, and abnormal thyroid function tests (most frequently hypothyroidism) [96,97]. Diagnosis of these immune-related adverse events should be made rapidly after excluding other etiological causes. Of particular interest is immune-mediated hepatitis, particularly with immunotherapy schemes such as nivolumab + ipilimumab or durvalumab + tremelimumab. Management and treatment with steroids should consider if the patient presents with cirrhosis, particularly with portal hypertension [96]. High-dose steroids in patients with cirrhosis and CSPH may induce further hepatic decompensation with ascites development. Thus, appropriate diagnoses and management require to exclude of other etiological causes, and a reduced dose of steroids from that recommended in other non-cirrhotic subjects. In a dual cohort study of patients treated with anti-PD-L1 comparing patients with HCC (72% with cirrhosis), and a non-HCC cohort, liver immune-related adverse events were more frequent, early in the treatment course, and more severe in patients with HCC [98]. However, complete resolution of these events was more frequently observed in the HCC cohort (72%vs 58%), even without steroid requirement (16%vs 75%). Consequently, an appropriate diagnostic approach should be undertaken before steroid indication, particularly in patients with cirrhosis. Rechallenge of immunotherapy may be considered after immune-related adverse events, if considering individual benefits and risks [99]
13Future perspectives: molecular and clinical decision-making processEmerging research highlights the molecular heterogeneity of HCC. Genomic analyses have identified distinct molecular subclasses broadly categorized into proliferative and non-proliferative phenotypes [100–102]. The proliferative subclass is characterized by activation of signaling pathways such as PI3K–AKT–mTOR, RAS–MAPK, and MET, frequently associated with TP53 mutations and higher AFP levels. In contrast, the non-proliferative subclass often involves WNT/β-catenin signaling and is typically associated with better differentiation and lower AFP levels [103]. In addition, immune profiling has identified distinct tumor immune microenvironment patterns, including immune-active, immune-excluded, and immune-desert phenotypes [104,105]. These patterns represent an area of active investigation [106–108]. Although a plausible explanation has been proposed of a reduced immune response in metabolic associated steatotic liver disease (MASLD) [109], this figure has not been observed in recent meta-analysis [110]. Future clinical trials may increasingly incorporate molecular and immunologic biomarkers to guide patient selection and optimize therapeutic strategies [111].
14ConclusionsThe interpretation of systemic therapy trials in hepatocellular carcinoma requires careful consideration of methodological, biological, and clinical factors. While overall survival remains the most robust endpoint in oncology trials, surrogate endpoints such as PFS and ORR must be interpreted cautiously in HCC due to the influence of competing risks related to underlying liver disease, selection bias (informative censoring), and misclassification errors (information bias). A comprehensive understanding of survival analysis, including Kaplan–Meier estimation, effect measures including HR, proportional hazards assumptions, and competing risk models, is essential for accurately evaluating treatment effects. Clinicians must recognize the heterogeneity of patient populations enrolled in clinical trials and the limitations of cross-trial comparisons.
In addition to improving survival outcomes, modern systemic therapies have introduced the possibility of conversion from “non-curative” to “curative” treatment strategies in selected patients with hepatocellular carcinoma. These developments challenge traditional stage-based treatment paradigms and emphasize the need for continuous reassessment of treatment response. Nevertheless, the term “cure” remains obscure after immunotherapy responses and should be appropriately individualized. As therapeutic strategies continue to evolve with the integration of immunotherapy, combination regimens, and personalized treatment approaches, multidisciplinary collaboration and methodological rigor will remain essential for translating clinical trial evidence into meaningful improvements in patient care.
Authors contributionsAll the authors approved the final version of the manuscript.
Funding statementThis research received no specific grant from any funding agency in the public, commercial, or non-profit sectors.
The authors of this manuscript have the following conflicts of interest to disclose:
- Federico Piñero, ORCID number 0000-0002-9528-2279: No direct conflicts of interest related to this research. Speaker honoraria, advisory board, and grants from ROCHE, BAYER, LKM Knight, RAFFOAstraZeneca. Has received grants from the National Cancer Institute (ID-19), and the Argentine National Institute of Medical Innovation (PICT 2017-1986).
- Arndt Vogel ORCID number 0000-0003-0560-5538
- Stephen L. Chan ORCID number 0000-0001-8998-5480. SLC serves as an advisory member for AstraZeneca, MSD, Eisai, BMS, Ipsen, and Hengrui, received research funds from MSD, Eisai, Ipsen, SIRTEX, and Zailab, and honoraria from AstraZeneca, Eisai, Roche, Ipsen, and MSD.
We would like to thank the Editor-in-Chief of Annals of Hepatology for this kind invitation to write this special article.












