The pilot policy on public data openness, a crucial initiative for advancing digital governance and leveraging data as a key production factor, has received limited systematic research regarding its mechanisms for enhancing urban innovation. Based on panel data from 282 prefecture-level cities in China spanning 2005–2023, this study employs a staggered difference-in-differences model to empirically identify the causal impact of the policy on urban innovation and uncover its underlying pathways. Notably, public data openness significantly enhances urban innovation, primarily through three channels: stimulating investment, fostering entrepreneurial activity, and attracting talent. Regional heterogeneity is evident, driven by differences in economic openness, the development of new quality productive forces, digital infrastructure, and fiscal expenditure on science and technology. This study enriches the theoretical literature on the relationship between public data openness and urban innovation by clarifying the unique institutionalized role of data as a production factor within urban innovation systems. Moreover, it offers actionable theoretical and methodological insights for policy evaluation through innovative research design and identification strategies, which can help policymakers effectively assess the impact of public data openness on urban innovation outcomes.
As the fifth major factor of production, data is a new and vital driver of growth and has evolved into a core resource powering high-quality economic development and shaping the digital economy. The openness of public data provides the foundation for optimizing the market-based allocation of data elements, playing a crucial role in unlocking their potential and strengthening the high-quality development of China’s economy. In 2023, China’s total data output reached 32.85 ZB, marking a year-on-year growth of 22.44%. Public data accounted for the largest share of this total and plays a fundamental, leading, and exemplary role in the development and utilization of data resources (Fang et al., 2023). Policy documents underscore that opening public data is a critical step in expanding the scale of data supply, improving supply efficiency, and enabling the circulation of data elements. These documents stress the relevance of dismantling institutional barriers to public data flows and striking a balance between efficiency and fairness. By doing so, the value of data as a production factor can be fully unlocked, fostering better decision-making and driving innovation across diverse sectors.
Contextually, the pilot policy for public data openness was introduced as a key initiative, unlocking the potential of data as a production factor and address circulation barriers. In 2012, China’s first public data openness platform was launched in Shanghai. Subsequently, in 2015, 2016, and 2018, China issued a series of policy documents, including the Action Outline for Promoting Big Data Development, Interim Measures for the Administration of Government Information Resource Sharing, and Work Plan for the Pilot Opening of Public Information Resources. These documents standardized public data openness across various dimensions, such as the digital economy, digital government, data governance, and data resources, providing a foundational guarantee for policy implementation and advancing the management and operation of public data. By July 2024, 243 provincial and municipal governments in China launched public data openness platforms, representing an increase of approximately 8% in the total number of platforms compared with 2023, thereby demonstrating a trend of continuous expansion (Fig. 1). This progress has established institutional and practical foundations for unleashing the value of data elements in urban innovation and economic development.
Public data inherently embodies a dual nature, characterized by data attributes and publicness (Shen & Feng, 2023). Therefore, public data is endowed with the unique potential to reshape urban innovation dynamics. The traceability, originality, and broad coverage of public data can significantly reduce information asymmetry and lower the marginal cost of knowledge acquisition. Moreover, the noncompetitive and nonexcludable nature of public data enables government entities, enterprises, and research institutions to access data resources extensively at a relatively low cost (Fang et al., 2023; Jones & Tonetti, 2020). Owing to the pronounced externality characteristics inherent in innovation activities, urban innovation systems have long been constrained by several key challenges, including insufficient investment in information, low efficiency of knowledge recombination, and inadequate government–enterprise interaction (Zhu & Zhang, 2024). Expanding the openness and accessibility of public data should alleviate these challenges, providing a stronger resource foundation and improved collaborative conditions for innovation.
However, a notable theoretical tension persists in the existing literature regarding the pivotal question of whether public data openness can genuinely enhance urban innovation capability. Although studies on open government data broadly affirm its role in improving transparency and public service efficiency, there is no scholarly consensus on its capacity to drive systemic, city-level innovation gains. This uncertainty primarily stems from three critical research gaps. First, existing research has predominantly focused on microlevel outcomes (Mergel et al., 2018; Shen & Lin, 2025), such as application development and entrepreneurship, with a notable lack of systematic investigation into the broader changes in the overall innovation capability of cities. Second, because public data is often correlated with unobservable structural characteristics of local governments—such as their digital governance capacity, level of economic development, and foundational innovation conditions—its use can cause endogeneity issues such as reverse causality and omitted variable bias. Current empirical studies lack effective identification strategies to address these concerns, thus failing to establish a reliable causal inference regarding the impact of public data openness on innovation. Finally, the specific mechanisms through which public data openness influences urban innovation—such as reducing information costs, promoting data-driven industrial expansion, or enhancing government–market collaboration—still lack systematic empirical examination.
China’s public data resources rank among the world’s foremost in terms of scale and coverage (Reinsel et al., 2018). Simultaneously, distinct bottlenecks persist within urban innovation systems (Cao et al., 2013), particularly concerning information integration, knowledge recombination, and government–enterprise coordination; this context makes the aforementioned research gaps particularly salient within the Chinese setting. Consequently, determining whether public data openness can genuinely enhance urban innovation capability and systematically examining its mechanisms has become a research agenda of academic relevance and practical urgency. To this end, this study leverages the staggered rollout of China’s public data openness pilot policy as a quasinatural experiment and employs a difference-in-differences (DID) approach to identify its causal effect on urban innovation. This design provides robust empirical evidence on the macrolevel innovation impact of public data openness and enables an effective investigation of its potential channels, including the reduction of information costs, stimulation of the digital economy, and alleviation of policy uncertainty. The study deepens the understanding of state-led data openness institutions as an emerging driver of urban innovation and offers policy implications for optimizing the governance of data elements and more fully unleashing the value of public data during China’s digital transformation.
Literature reviewMicro and macro effects of public data opennessAs the pilot policy for public data openness has been progressively implemented across Chinese cities, the academic community has started focusing on its practical effects, conducting evaluations from micro and macro perspectives. At the microlevel, public data openness can reduce risks in private lending (Wang et al., 2025), enhance firms’ digital innovation output (Li et al., 2025), improve the quality of corporate services (Ahmadi Zeleti et al., 2016), stimulate innovation in private enterprises (Li et al., 2025), increase corporate investment (Wang et al., 2024; Ye et al., 2025), advance the development of the financial sector (Ouyang et al., 2024), and assist enterprises in cultivating new quality productive forces (Dong et al., 2025). Furthermore, public data openness facilitates cross-regional capital mobility, attract new firm entry, and drive corporate upgrading (Lan et al., 2024; Zheng & He, 2025; Zou et al., 2025).
At the macrolevel, public data openness policies can promote coordinated regional development (Fang et al., 2023), enhance fiscal transparency (Liu et al., 2025), increase employment (Liu et al., 2024), foster urban industrial integration (Guo et al., 2026), and improve urban entrepreneurial vitality while optimizing the business environment (Cai et al., 2024; Ding & Bai, 2025). While these studies are helpful for understanding the policy effects of public data openness, they have predominantly concentrated on its immediate outcomes, leaving a gap in the systematic and in-depth analysis of its long-term effects on innovation.
Measuring urban innovation levels and influencing factorsResearch on urban innovation levels initially focused on measurement methods. Scholars have approached this focus from the perspectives of innovation inputs (Liu et al., 2021) and innovation outputs (Wu & Zhang, 2021). Some have developed multi-indicator evaluation systems that encompass innovation potential, outputs, and inputs to measure urban innovation (Jiao et al., 2017), although these primarily concentrated on the provincial level or select economically developed regions. At present, using patent and related data to gauge urban innovation has become the mainstream approach (Gu et al., 2025; Wang & Deng, 2022; Wang et al., 2025; Wen et al., 2023; Zhang et al., 2022). Therefore, scholars have further explored influencing factors, discovering that urban innovation levels are affected by government-guided funds (Zhuang et al., 2025), the comprehensive development level of interprovincial internet infrastructure (Han et al., 2019), financial technology (Li et al., 2024), urban amenities (Bai & Wei, 2024), and the digital economy (Li & Wu, 2024; Zhang et al., 2023; Zhao et al., 2024). Most of these factors are directly or indirectly policy-related, providing a theoretical foundation for studying the impact of policies on urban innovation.
Relationship between public data openness and urban innovation: research gapsAlthough substantial research examines public data openness and urban innovation separately, their systematic relationship remains insufficiently clarified. Existing studies predominantly emphasize the positive effects of open government data on innovation proposals, transparency, and social value creation (Pedersen, 2020). However, a key theoretical tension persists. On the one hand, public data can reduce information asymmetry and lower transaction costs, thereby stimulating innovative activity. On the other hand, public data may exacerbate digital divides, reinforce resource concentration, and widen intercity disparities (Zheng & He, 2025). This dualistic effect raises an unresolved question: does public data openness ultimately enhance systemic urban innovation capability or generate uneven outcomes across cities? Building on this tension, this study investigates the overall impact of public data openness pilot policies on urban innovation performance, considering specific institutional and regional contexts.
Beyond the aggregate effect, another tension arises between microlevel innovation gains and systemic-level transformation. Prior research often focused on firm-level applications of open data, such as product development or service improvement, yet provided limited insight into how policy-driven data openness reshaped urban innovation systems as a whole. The literature suggests several potential channels, including improved capital allocation, data-driven industrial expansion, and strengthened government–market interaction (Goldfarb & Tucker, 2019; Wirtz et al., 2022), but lacks systematic empirical validation of these mechanisms. Therefore, this study explicitly examines three key pathways through which public data openness may influence urban innovation: the reallocation of investment resources, stimulation of entrepreneurial activity, and agglomeration of human capital. These mechanisms correspond to capital input, market dynamism, and knowledge accumulation within the urban innovation system.
Theoretical frameworkImpact of public data openness on urban innovation levelsAgainst the backdrop of the accelerated evolution of the digital economy, data has transformed from an auxiliary informational resource into a key factor that profoundly shapes resource allocation and value creation. Compared with traditional production factors such as capital and labor, data exhibits distinctive characteristics, including nonrivalry, reusability, and low marginal cost (Jones & Tonetti, 2020). The value of data accumulates through sharing, integration, and continuous reuse and generates significant spillover and network effects as it flows across actors (Jones & Tonetti, 2020). Meanwhile, the accessibility, usability, and reliability of data do not emerge naturally but heavily depend on institutional supply and governance arrangements (Kitchin, 2014). Concerning public data, in particular, its production, integration, and disclosure are embedded within government-led regulatory frameworks and organizational structures, highlighting its pronounced institutional attributes (Wirtz et al., 2022). Therefore, public data should be understood as a technical resource and an institutionalized production factor defined and allocated through policy design, institutional arrangements, and governance mechanisms.
Mainstream innovation theories have traditionally focused on endogenous technological progress or the reconfiguration of conventional inputs while paying limited attention to the institutionally embedded nature of data. As policies on public data openness continue to advance, the mechanisms of data provision and the governance structures surrounding them have emerged as critical conditions shaping innovation performance (Wirtz et al., 2022). Treating data merely as an additional input within an extended factor framework is insufficient to explain how public data openness reshapes resource allocation, activates market actors, and transforms the innovation ecosystem through institutionalized supply mechanisms. Therefore, reconstructing the theoretical framework by systematically conceptualizing public data as an institutionalized production factor and clarifying its underlying mechanisms of influence is essential.
The implementation of public data openness policies can promote the sharing of data elements and fully leverage their potential value, playing a substantial role in enhancing the innovation level of cities. On the one hand, the wide coverage and large data volume characteristics of public data can reduce the cost of information search and acquisition (Goldfarb & Tucker, 2019) and enable innovation entities to explore the potential value of data elements more deeply (Vetrò et al., 2016). On the other hand, public data openness can provide a favorable external environment for the flow of data elements and knowledge spillover, generating a multiplier effect (Chen & Xu, 2024). Based on the Cobb–Douglas production function model (Cobb & Douglas, 1928), we identify data elements as a crucial production factor for enterprises and construct the theoretical model.
Suppose an enterprise invests three factors, namely, capital (K), labor (L), and public data (D), in its production process. The marginal costs of these three factors are, respectively, rK, rL, and rD/(1+Z), generating innovative output (Y). An enterprise’s innovation output (Y) can represent its innovation level. The degree of public data openness in this region is represented by Z, with (0, 1] range. Owing to the public nature of the public data itself, coupled with the fact that enterprises will collect their data, Z will not be equal to 0. However, if the region implements a public data openness policy, Z can equate to 1. The cost for enterprises to obtain data elements is related to whether the local area implements a public data openness policy. If no such policy is implemented, the cost for enterprises to obtain data is rD. If yes, it becomes easier for enterprises to obtain data, and the cost is lower. Subsequently, the cost of data elements is rD/(1+λZ). To simplify the derivation, we take λ=1, and the cost of data elements becomes rD/(1+Z). The efficiency of the production of an enterprise is influenced by the technological level (A), reflecting the accumulation of production technology and knowledge and is related to the externality of data elements and the diffusion of technology, specifically A=A0+δZ. Subsequently, the innovation production function of the enterprise can be determined as follows:
According to microeconomic theory, the profit function of enterprises is as follows:
To analyze the optimal usage of enterprise data elements, we take the first-order conditions of the enterprise profit function as follows:
The optimal input of enterprise data elements can be obtained as follows:
To analyze the impact of public data openness policies on the usage of enterprise data elements, we take the derivative of D* with respect to Z and obtain the following:
Subsequently, the public data openness policy will enhance the optimal usage of enterprise data elements.
To analyze the impact of public data openness on the innovation output of enterprises, we first present the complete production function. Incorporating the technological level, the production function is expressed as follows:
To analyze the impact of public data openness policies on the innovation level of enterprises, we derived the innovation output with respect to these policies and obtained the following equation:
If the derivative of innovation output with respect to public data openness policy is consistently positive, then such a policy exerts a positive impact on the innovation level of enterprises.
We believe that the innovation level of a city can be approximately expressed by the sum of the innovation levels of enterprises as follows:
Therefore, when policy Z = 1, the innovation output of each enterprise increases, leading to a corresponding rise in the city’s overall innovation output and an enhancement of its innovation level. Furthermore, the positive externalities of data elements will further amplify this effect. In conclusion, the implementation of public data openness policies will enhance the innovation level of cities. Building on this conclusion, this study presents the following hypotheses.
H1 The pilot policy for public data openness can enhance the innovation level of cities.
Investment level, entrepreneurial activity, and human resource concentration constitute the primary channels through which public data openness affects urban innovation. These three dimensions are selected because they correspond to the core innovation resources of capital input, market dynamism, and human capital, forming the foundational pillars of urban innovation systems (Cooke et al., 1997). Investment captures the scale and efficiency of financial resource allocation and reflects the capital foundation of innovation activities (Hall & Lerner, 2010). Entrepreneurial activity represents the vitality of market actors and the creation of new organizational entities, serving as a direct indicator of opportunity recognition and innovative experimentation (Shane & Venkataraman, 2000). Human resource concentration reflects the accumulation and mobility of skilled labor, underpinning knowledge creation, diffusion, and recombination (Moretti, 2004). Compared with other potential channels, these three mechanisms provide a more systematic and theoretically grounded representation of how public data, conceptualized as an institutionalized production factor, operates through resource allocation, firm formation, and knowledge generation.
Investment level typically refers to the scale of infrastructure capital allocated by a city to fields such as research and development, technology, and industrial development, along with the efficiency of its utilization. In this context, the public data openness pilot policy indirectly enhances urban innovation capability by improving the investment environment. First, the pilot policy can expand the available pool of data resources (Fang et al., 2023), mitigate information asymmetry, reduce institutional transaction costs, and improve the matching efficiency of information resources (Long et al., 2025). This reduces investment uncertainty and promotes capital agglomeration. Second, the pilot policy aids governments in optimizing administrative functions, enhancing transparency, and accelerating the development of new industrial parks and innovation R&D centers (Li & Liu, 2022). Higher policy transparency and improved digital infrastructure can attract technology- and innovation-driven enterprises to locate in pilot cities, further concentrating investment. The resulting increase in the investment level provides solid capital support for urban innovation, promoting the enhancement of the city’s overall innovation capability. Therefore, the following hypothesis is proposed:
H2 The public data openness pilot policy influences urban innovation through the investment level pathway.
Entrepreneurial activity is typically measured by the number of newly established enterprises in a city. In this context, the public data openness pilot policy indirectly promotes urban innovation levels by enhancing entrepreneurial activity. First, the pilot policy can reduce information asymmetry (Huang & Yu, 2023) and the uncertainty associated with corporate innovation (Cai & Ma, 2021; Shen & Lin, 2025), enabling entrepreneurs to identify and seize market opportunities more rapidly, thereby increasing the number of new firm formations. Second, public data openness platforms unleash data dividends, lowering the costs of R&D and market exploration, further stimulating entrepreneurial vitality. The rise in entrepreneurial activity attracts more high-quality innovation resources to the city and fosters a more vibrant culture of innovation, facilitating the translation and diffusion of innovative outcomes. This ultimately enhances the overall innovation level of the city. Considerably, the following hypothesis is proposed:
H3 The public data openness pilot policy influences urban innovation through the entrepreneurial activity pathway.
Human resource concentration refers to the spatial clustering of highly skilled and innovative talent within a specific region. Contextually, the public data openness pilot policy indirectly enhances urban innovation levels by fostering a human resource concentration effect. First, the development of public data openness platforms facilitates the cultivation of emerging digital industrial ecosystems, stimulating employment and attracting a flow of professional talent (Huang & Yu, 2023), augmenting the city reservoir of innovative human capital. Second, pilot cities for public data openness typically possess a more open, equitable, and liberal social climate (Du, 2020), providing a conducive environment for career development and innovative practices of highly skilled individuals. Talent concentration fosters the aggregation of specialized knowledge and innovative resources and further drives the enhancement of urban innovation capacity by unlocking the latent value of data and harnessing its dividends as a production factor. Thus, the following hypothesis is proposed:
H4 The public data openness pilot policy influences urban innovation through the human resource concentration pathway.
Fig. 2 shows the theoretical framework of this study.
Methodology and materialsMethodologyBenchmark regression model specificationBai et al. (2022) and Zhang et al. (2023) treated the launch of public data openness platforms in various cities as a quasinatural experiment and employed a staggered DID model to assess its impact on urban innovation levels. Specifically, cities with launched platforms are designated as the treatment group, whereas those without served as the control group. By controlling for city and time fixed effects, the model analyzes differences in urban innovation levels before and after policy implementation to identify the causal effect of the policy shock. Additionally, this study incorporates robustness tests, including parallel trend tests, dynamic effect analysis, Propensity Score Matching combined with DID (PSM–DID), and instrumental variable methods, to ensure the reliability of the findings. The research logic can be summarized as follows: the launch of public data openness platforms induces changes in investment, entrepreneurship, and talent resources, which drives improvements in urban innovation levels and systematically examines the policy’s effect and its transmission mechanisms.
The staggered DID model is adopted for three primary reasons. First, the launch timing of public data openness platforms exhibits significant variation across cities. The staggered nature of this policy shock, with its exogenous and phased characteristics over time, provides the necessary identification conditions for the DID approach. Second, by comparing the treatment group with the control group while controlling for city and time fixed effects, the method can effectively mitigate endogeneity concerns related to the policy, such as unobservable differences in cities’ digital governance capacity, economic foundations, and innovation environments. Third, compared with simple cross-sectional or before-after analyses, the DID method more clearly captures dynamic changes before and after the policy shock. The DID method addresses issues arising from staggered treatment timing and heterogeneous treatment effects through multiple robustness checks (e.g., PSM–DID and Goodman–Bacon decomposition), enhancing the reliability of causal inference. Furthermore, the model design accounts for potential contemporaneous policy interference and measurement errors in innovation, ensuring the robustness and interpretability of the results.
In summary, building on and fully leveraging existing research, this study adopts the staggered DID model and systematically evaluates its underlying assumptions and potential limitations throughout the analysis. This approach aims to ensure comprehensive research design and prudent inference, thereby guaranteeing the robustness and reliability of the empirical findings. Consequently, cities that have launched public data openness platforms are designated as the treatment group, while the remaining cities constitute the control group. The staggered DID model is employed for empirical testing, with the specific regression model specified as follows:
The subscript i represents the city and t represents the year. Ur_Innovationit represents the innovation level of City i in year t, Openit represents the dummy variable that goes online on the public data openness platform. If city i goes online on the public data openness platform in year t, it takes 1; otherwise, it takes 0. Control_Varit is the set of control variables, δi represents the urban fixed effect, δt represents the fixed effect of the year, εit is a random disturbance term. Considering that the random perturbation terms of samples within the same city may be potentially correlated, this study clusters the standard misclustering at the city level. α is the constant term, and β is the regression coefficient; overall, they indicate the impact of public data openness on urban innovation in China.
Parallel trend testIn the context of this study, employing a staggered DID model requires satisfying the parallel trend assumption. This assumption posits that prior to the policy intervention, the treatment and control groups follow a common trend. In other words, the change in urban innovation levels for cities without public data openness platforms (the control group) represents a valid counterfactual for what the change would have been for cities with such platforms (the treatment group) had they not launched them. To mitigate potential bias from the uneven distribution of treated samples in practice, this study follows the approach of Bai et al. (2022) and Zheng & He (2025) by employing an event study framework. Specifically, relative periods earlier than −4 are recoded as period −4, and period −1 is designated as the base period. The estimation model is specified as follows:
TreatN represents the dummy variable of the “event” of going online on the public data openness platform, and the time dummy variable N represents the observed values of each city in the n years before, in the current year, and in the n years after it was established as a pilot city. Specifically, when N = −1, it denotes the year prior to the launch of the public data openness platform; in other words, Treat−1=1. When N = 1, it denotes the first year of the platform going online; in other words, Treat1=1. The dummy variables for other nonpilot cities are all 0.
Mechanism testBuilding on the preceding theoretical analysis, this study argues that the public data openness pilot policy enhances urban innovation primarily by increasing investment, fostering entrepreneurial activity, and promoting human resource concentration. To empirically examine these mechanisms through which the policy influences innovation in Chinese cities, the study adopts Jiang’s (2022) methodological framework for analyzing mechanisms and transmission channels and constructs the following mediation effect model:
Betw_Varit represents the mechanism variable, which is used, respectively, used to capture the investment level (Invest_level), entrepreneurial activity level (Entr_activity), and agglomeration effect of human resources (Hu_capital). Other variables are consistent with the previous text. If coefficients θ and ω are significant and in line with expectations, this result indicates that the mechanism variable holds true and the mechanism test is verified.
Variable definitionsThe explained variable of this article is the urban innovation level (Ur_Innovation), the explanatory variable is the public data openness policy (Open), and the control variables include the economic development level, the added value of the secondary industry, the total urban population, the urbanization rate, and the higher education level, totaling five.
- (1)
Urban innovation level (Ur_Innovation). Referring to the studies of Jin et al. (2019); Li & Zheng (2016); Zhao et al. (2020), and Wang et al. (2023), the level of urban innovation was measured by the number of patent authorizations from the perspective of innovation output. Compared with a patent application, the patent authorization review process is more precise and better represents the verified and effective level of innovation output.
- (2)
Public data openness policy (Open). Referring to the research of Liu et al. (2024); Pan et al. (2023); and Zhao & Hao (2023), the pilot policy of public data openness was regarded as a quasinatural experiment, and the interaction term (Group×Post) between the dummy variable of city type and the dummy variable of policy implementation time was used to represent the policy processing effect of the public data openness policy. Specifically, in this study, the pilot city Group for the open launch of public data is set as 1 as the experimental group, and the remaining cities are set as 0 as the control group. Set the time dummy variable Post to 0 and 1, respectively, before and after the implementation of the pilot policy. Since the launch of urban public data openness platforms occurred in batches from 2012 to 2023, the time dummy variables for different cities with open public data are not aligned.
- (3)
Control variable (Control_Var). The selection was made per Zheng & He’s (2025) and Bai et al.’s (2022) studies. Specifically, the level of economic development (lnGDP) is measured by the natural logarithm of the gross domestic product. The added value of the secondary industry (lnAdd) is measured as the natural logarithm of the added value of the secondary industry. The total urban population (lnpeople) is expressed as the logarithm of the urban registered population. Following Tang et al. (2022), the urbanization rate (ubr) is primarily measured by U/(U+R), where U and R denote urban and rural permanent resident populations. Respectively, as a robustness check, we also employ N/P, the ratio of nonagricultural population (N) to year-end total population (P). The level of higher education (edu) is expressed by the number of regular institutions of higher learning.
This study selects 282 prefecture-level cities in China from 2005 to 2023 as the research sample. Cities established after 2011 owing to administrative division adjustments, such as Bijie City and Tongren City in Guizhou Province, and those with severe data deficiencies such as Lhasa City, Danzhou City, and Sansha City, are excluded to construct a balanced panel dataset to the greatest extent possible. The launch dates of local governments’ public data openness platforms are identified by referencing the China Local Government Data Openness Report published by the Digital and Mobile Governance Laboratory of Fudan University and existing literature (Liu et al., 2024; Pan et al., 2023; Zhao & Hao, 2023).
City-level statistical data are sourced from the China City Statistical Yearbook and China National Intellectual Property Administration. Missing values for individual samples are addressed by consulting local statistical yearbooks and bulletins or by applying linear interpolation.
Variable definitions and descriptive statistics are presented in Table 1.
Variable definitions and descriptive statistics.
Table 2 reports the regression results of the impact of the public data openness policy on urban innovation levels. Column (1) presents the estimated effect of the public data openness policy on urban innovation, controlling for city and year fixed effects. The coefficient for Open is positive and statistically significant at 1%. Columns (2) through (6) sequentially add control variables to the specification in Column (1). In all specifications, the coefficient for Open remains significantly positive at 1%, consistently indicating a positive relationship between the public data openness policy and urban innovation levels, supporting Hypothesis H1. Theoretically, driven by the public data openness policy, pilot cities can significantly improve information acquisition environments and reduce institutional transaction costs, enabling capital, entrepreneurial activities, and highly skilled talent to flow more efficiently into innovation-intensive sectors. Specifically, the policy enhances the precision of investment decisions and the feasibility of entrepreneurial projects and strengthens the agglomeration effects of human capital and specialized knowledge, thereby boosting overall urban innovation output.
Baseline regression results.
Note: Values in parentheses represent robust standard error, with “***” and “**” denoting significance at 1% and 5%, respectively.
Fig. 3 depicts a series of estimated coefficients of TreatN within a 95% confidence interval, reflecting the dynamic effect of public data openness policies on the innovation level of Chinese cities.
Results demonstrate that the estimated coefficients of TreatN before the policy implementation are insignificant and small in magnitude, indicating no significant difference in the changing trend of urban innovation levels between Chinese cities in the treatment group and those in the control group. However, after the launch of the public data openness platform, the estimated coefficient of TreatN was significantly positive, indicating that there were significant differences in the changing trends of urban innovation levels between the experimental group and the control group after policy intervention, and the parallel trend test was confirmed. Moreover, the estimated coefficient of TreatN is constantly increasing, indicating that the impact of the pilot policy of the public data openness platform on the innovation level of Chinese cities is gradually strengthening; in other words, the promoting effect is gradually being released.
Robustness testPlacebo testAlthough this study has controlled for various urban characteristic variables within the quasinatural experiment framework, the observed correlation between the public data openness policy and urban innovation levels could be spurious because of unobserved urban characteristics. Per Fang et al. (2023); Zheng & He (2025), and Bai et al. (2022), a placebo test is conducted to verify that the findings are driven by the actual launch of public data openness platforms rather than random chance. Specifically, using Stata, a pseudotreatment group dummy variable Grouprandom and a pseudopolicy shock dummy variable Postrandom are randomly generated for the 282 sample cities. A fictional public data openness pilot policy is then simulated through 500 random shocks. In each iteration, 195 cities are randomly selected as the treatment group, and the policy timing is randomly assigned. Fig. 4 shows the kernel density distribution of the estimated coefficients and their corresponding p-values from these placebo tests. The coefficients from the randomized policies are centered around 0 and are significantly smaller in magnitude than the true estimated coefficient (0.0904). Furthermore, the impact of the public data openness policy on China’s urban innovation levels is not confounded by unobserved city characteristics, confirming the validity of the quantitative assessment.
Multitime point PSM–DID modelTo alleviate the endogeneity problem caused by sample selectivity bias, this study refers to the studies of Yu et al. (2022) and Liang et al. (2023) and adopts the approach of 1:1 matching with a nearest neighbor caliper without putting it back. Specifically, in this study, all control variables (Xia et al., 2024) are used as covariates for matching, and the experimental groups affected by policies are taken as dependent variables to conduct “one-to-one, no return” nearest neighbor matching. Fig. 5 shows that neither the results before nor after matching can reject the null hypothesis that there is no systematic difference between the treatment group and the control group. Fig. 6 shows that the treatment group and the control group have a sufficiently substantial overlap within the propensity score range, satisfying the common support assumption. Fig. 7 shows that after matching, the distribution of the two groups’ samples is more similar, suggesting that PSM effectively improves the comparability of the samples. The matching estimation results in column (1) of Table 4 show that the regression coefficient of Open is significantly positive at 1%, further indicating that the conclusion of this study is robust.
Goodman–Bacon decompositionDuring the estimation process of the staggered DID model, the two-way fixed effects (TWFE) estimator is equivalent to a weighted average of all possible two-period DID estimators within the sample. This approach can potentially lead to issues of nonrobustness or even negative weights owing to heterogeneous treatment effects (Goodman-Bacon, 2021). Per Baker et al. (2022) and Dai & Zhao (2024), this study conducts a Bacon decomposition. This procedure calculates the coefficient and weight for each distinct DID comparison group to examine the potential bias in the TWFE staggered DID estimate. From Fig. 8, the proportion of “bad treatment groups” is small, indicating that the impact of the public data openness pilot policy on China’s urban innovation levels is not driven by this type of comparison. Table 3 demonstrates that the overall DID estimate herein primarily stems from comparisons where the control group comprises units that were never treated, accounting for a high weight of 71.1%. By contrast, comparisons that could introduce bias—those using units treated earlier as controls for units treated later—carry a weight of only 5.3%, and their DID estimates are negative. In conclusion, the core findings of this study are relatively robust, and the staggered DID estimation results are reliable.
Re-examination based on sample screening(I) Winsorization. To mitigate extreme value influences, we re-estimate models after winsorizing urban innovation levels at the 1st and 99th percentiles. As column (2) of Table 4 indicates, the policy coefficient remains positive and statistically significant at 1%, confirming that public data openness robustly promotes urban innovation—consistent with baseline findings.
Robustness check estimation results.
Note: Values in parentheses represent robust standard error, with “***,” “**,” and “*” denoting significance at 1%, 5%, and 10%, respectively.
(II) Exclusion of Direct-Controlled Municipalities. Addressing potential advantages in economic development, resource access, policy bias, and geographic positioning, we exclude Beijing, Tianjin, Shanghai, and Chongqing from the sample. The results in column (3) of Table 4 demonstrate that the policy coefficient persists positively significant at 1%, reaffirming the core conclusion’s robustness.
Exclude the influence of other policiesWithin the research time range of this article, the national-level pilot policies for creating entrepreneurial cities introduced in 2010 and the pilot policies for establishing smart cities in batches in May 2012, 2013, and 2014 have an impact on the innovation level of cities. Referring to the research approach of Bai et al. (2022), to eliminate the interference of other policies, dummy variables for the national-level pilot policies for creating entrepreneurial cities (Ac_Policy) and smart cities (Sma_Policy) were generated and successively added to the benchmark regression model. Results are shown in Columns (4) and (5) of Table 4. With respect to the smart city pilot policies, since some prefecture-level cities only select a certain district or county within the city as a pilot when establishing a smart city, choosing that prefecture-level city as a pilot city would underestimate the impact of smart cities. Therefore, this article excludes such prefecture-level cities from the sample. When other policies are successively added to the pilot policies for public data openness, they are still significantly positive at 1%, which proves that the results of this study are robust.
Alternative variable measurementsGiven that the dependent variable is urban innovation level, initially measured from the perspective of innovation output using the number of patents granted, this study adopts alternative measurement approaches to ensure comprehensiveness. Specifically, the dependent variable is replaced with six different metrics: the number of invention applications (ZZ1), utility model applications (ZZ2), design applications (ZZ3), invention grants (ZZ4), utility model grants (ZZ5), and design grants (ZZ6). These variables collectively encompass various forms of innovation, including social, service, and digital innovation. The regression results using these alternative measures are presented in Table 5. The public data openness pilot policy continues to demonstrate a positive and significant impact on urban innovation levels from these different perspectives, confirming the robustness of the main findings. The policy exhibits a particularly strong effect on utility model patents, with an impact coefficient reaching 0.4031.
Estimation results with alternative variable measurements.
Note: Values in parentheses represent robust standard error, with “***” and “**” denoting significance at 1% and 5%, respectively.
Per Li & Liu (2022), this study measures the investment level as the ratio of fixed asset investment to urban area. From Column (1) of Table 6, the regression coefficient for Open is 0.6149 and is statistically significant at 1%, indicating that the public data openness pilot policy significantly enhances the agglomeration effect of urban investment, supporting Hypothesis H2. In other words, with the effective implementation of the pilot policy, the increase in the investment level provides more substantial financial support for urban innovation. Simultaneously, this increase in investment attracts a large number of technology- and innovation-driven enterprises to cluster in the pilot cities, generating an agglomeration momentum akin to a “flying geese” effect, which drives the overall improvement of urban innovation levels.
Mechanism test results.
Note: Values in parentheses represent robust standard error, with “***,” “**,” and “*” denoting significance at 1%, 5%, and 10%, respectively.
Per Huang & Yu (2023), this study uses the natural logarithm of the number of newly registered enterprises in the information transmission, computer services, and software industry—obtained from the Qichacha database—as a proxy for entrepreneurial activity. From Column (2) of Table 6, the coefficient for Openis positive and statistically significant at 1%, indicating that the public data openness pilot policy significantly promotes urban entrepreneurial activity, thus supporting Hypothesis H3. Compared with the benchmark regression coefficient of 0.0904, this coefficient is larger, suggesting that policy implementation attracts a greater number of new ventures in the information and data industries to locate in pilot cities. These enterprises drive the upgrading of digital infrastructure (Li et al., 2024) and provide technological and talent support for urban innovation. Furthermore, the increase in entrepreneurial activity strengthens the agglomeration effects of innovation factors, fostering positive interactions among investment, talent, and technology, thereby further elevating the overall level of urban innovation.
Human resource concentration effect (Hu_capital)Per Zou (2025), this study measures human resource concentration as the natural logarithm of the ratio of full-time undergraduate and college students in a prefecture-level city to its year-end total population. From Column (3) of Table 6, the coefficient for Open is positive and statistically significant at 5%, indicating that the public data openness pilot policy significantly promotes the concentration of human resources in cities, thereby supporting the pathway outlined in Hypothesis H4. Although the coefficient is only 0.0472, highly skilled talent is a core driver of innovation, and its agglomeration effect has a long-term and profound impact on urban innovation levels. Specifically, public data openness provides talent with richer innovation information and development opportunities and attracts and retains technical talent by improving entrepreneurial and research environments. This provides crucial human capital support for urban innovation and enhances the efficiency with which investment and entrepreneurial mechanisms translate into innovation output.
Heterogeneity analysisHeterogeneity by economic openness levelThe effectiveness of the public data openness policy may be constrained to some extent by the level of economic openness, as a city’s administrative rank often correlates with its economic autonomy and developmental capacity. Per Dong & He (2025); Zhao et al. (2020), and Kang et al. (2025), this study divides the sample into two groups based on whether a city is a provincial capital, a city specifically designated in the state plan, or a municipality directly under the central government. Cities falling into these categories are classified as having high economic openness, whereas others as having low economic openness. The regression results for this heterogeneity analysis are presented in Columns (1) and (2) of Table 7. Notably, the coefficients for both groups are statistically significant at 1%. However, the coefficient for cities with high economic openness is larger, reaching 0.3935, than those with low economic openness. The underlying reason may be that compared with ordinary prefecture-level cities, provincial capitals, cities specifically designated in the state plan, and municipalities possess greater administrative advantages and institutional convenience. These cities can allocate more talent, funding, and policy support to the launch and operation of public data platforms and enjoy greater economic autonomy. Consequently, the public data openness pilot policy exhibits a more pronounced effect on enhancing innovation levels in these high-economic-openness cities. Simultaneously, such cities are more likely to generate agglomeration effects of innovation factors, ultimately amplifying the policy’s role in promoting urban innovation.
Heterogeneity analysis results by economic openness level and level of new quality productive forces.
Note: Values in parentheses represent robust standard error, with “***” denoting significance at 1%.
New quality productive forces represent an advanced form of productive capacity aligned with the new development philosophy. These productive forces have given rise to novel cooperative models, such as the platform economy and the sharing economy, effectively enhancing urban innovation levels. Drawing on the research of Han et al. (2024), this study employs the entropy method to measure the level of new quality productive forces in 282 prefecture-level cities in 2024, focusing on three dimensions: new quality labor, new quality objects of labor, and new quality means of labor. The research sample is then divided into high-new-quality-productive-forces and low-new-quality-productive-forces regions based on the median value for heterogeneity analysis. The corresponding regression results are reported in Columns (3) and (4) of Table 7. Results suggest that compared with low-level regions, the coefficient for high-level regions is significantly positive at 1% and larger in magnitude. By contrast, the coefficient for the low-level group is not statistically significant. A possible explanation is that high-level regions are continuously transforming their growth drivers, gradually moving away from traditional economic growth patterns and development paths for productive forces. This transformation entails qualitative leaps in laborers, means of labor, and objects of labor within the agglomeration of production factors (Tan, 2025). By contrast, low-level regions experience slower agglomeration of newer and more complex production factors—such as digital labor and digital technologies—resulting in lagging development of emerging business forms. Therefore, the promotional effect of the public data openness pilot policy on urban innovation levels is more pronounced in regions with high levels of new quality productive forces.
Heterogeneity in digital infrastructureDigital infrastructure facilitates information transmission, R&D collaboration, and knowledge spillovers (Lan et al., 2024), contributing to urban innovation. However, significant disparities exist in the level of digital infrastructure development across different regions. Per Wang et al. (2023), this study constructs a comprehensive evaluation index system to measure digital infrastructure, covering input and output dimensions. For regions with partial data missing, the study draws on Chao et al.’s (2021) methods, using the proportion of keywords related to new digital infrastructure in local government work reports as a proxy for imputation. The sample is then divided into high- and low-digital-infrastructure groups based on the median level of digital infrastructure for heterogeneity analysis. Results are presented in Columns (1) and (2) of Table 8. The impact of the public data openness pilot policy on urban innovation levels is statistically significant in the high and low groups. However, the regression coefficient for the high-digital-infrastructure group is five times larger than that of the low-digital-infrastructure group. In other words, the effect of the policy on enhancing urban innovation is more pronounced and exhibits a clear multiplier effect in regions with advanced digital infrastructure. A possible explanation is that digital infrastructure leverages information technology to provide more efficient channels for the dissemination of the latest R&D achievements and cutting-edge knowledge spillovers (Pei & Zhang, 2025). Furthermore, the process of data openness and sharing is faster and more efficient in high-digital-infrastructure regions. This faster and more efficient process fosters the flow and allocation of data elements among various innovation actors, unlocking the latent value of public data (Ma & Cui, 2024). Consequently, the public data openness pilot policy demonstrates a more substantial effect in regions with well-developed digital infrastructure.
Heterogeneity analysis results by digital infrastructure level and fiscal expenditure on science and technology.
Note: Values in parentheses represent robust standard error, with “***” and “*” denoting significance at 1% and 10%, respectively.
Substantial fiscal expenditure on science and technology paves the way for urban innovation and equips cities with a highly skilled scientific and technological workforce. Consequently, the effect of the public data openness pilot policy on urban innovation may vary depending on the intensity of such fiscal expenditure. Per Li (2024) and Zhang & Luo (2024), this study measures fiscal expenditure on science and technology as the ratio of a prefecture-level city government’s spending on science and technology to its general budgetary expenditure in 2024. The sample is then divided into high and low fiscal expenditure groups based on the median value for heterogeneity analysis. Results are presented in Columns (3) and (4) of Table 8. The coefficients for both groups are positive and significant at 1%. However, the coefficient for the high-expenditure group is 0.1296, which is substantially larger than the coefficient of 0.0088 for the low-expenditure group; this indicates that the policy has a more pronounced effect on enhancing innovation in cities with higher fiscal expenditure on science and technology. The underlying reason may be that higher fiscal spending can swiftly attract the agglomeration of innovation factors such as innovative capital and talent. Moreover, the efficient concentration of these factors can amplify the positive externalities of knowledge spillovers Meng & Wu (2025), further leveraging the role of the public data openness pilot policy in boosting urban innovation levels.
Conclusion and DiscussionConclusionBased on panel data from 282 prefecture-level cities in China spanning 2005–2023, this study employs a staggered DID model to empirically examine the impact of the public data openness pilot policy on urban innovation levels. The policy significantly enhances urban innovation, a conclusion that remains consistent across multiple robustness checks, indicating a reliable positive effect of the policy on fostering innovation. Furthermore, the policy operates through multiple channels, primarily by increasing the investment level, boosting entrepreneurial activity, and concentrating human resources. Thus, the policy directly influences innovation output and indirectly promotes urban innovation by improving key aspects of the innovation ecosystem. Simultaneously, the effects of the policy exhibit significant heterogeneity. Factors such as a region’s degree of economic openness, level of new quality productive forces, digital infrastructure development, and scale of fiscal expenditure on science and technology contribute to variations in the policy’s impact on urban innovation across different regions. In conclusion, the public data openness pilot policy has an overall positive effect on enhancing urban innovation levels; however, its effectiveness is moderated by regional conditions and resource endowments. This underscores the relevance of tailoring policy design to account for regional disparities, aiming for more targeted innovation incentives.
ImplicationsTheoretical implicationsThe public data openness pilot policy enhances urban innovation through three inter-related mechanisms: higher investment, stronger entrepreneurial activity, and greater human resource concentration. Theoretically, these findings extend microlevel studies on data openness by embedding them in the broader framework of urban innovation systems and digital economy development (Harrison et al., 2012; Mergel et al., 2018). Public data is thus positioned as a production factor that operates through systemic channels shaping factor allocation at the urban level. First, public data openness reduces information asymmetry and transaction costs. Public data openness improves the investment environment and directs capital toward innovation-intensive sectors. This reinforces the link between transparency and capital allocation efficiency (Huang & Yu, 2023). Second, accessible data lowers uncertainty in opportunity recognition and experimentation. Accessible data stimulates entrepreneurship and clarifies its role as a transmission channel from institutional data supply to innovation outcomes (Cai & Ma, 2021; Chan, 2013). Third, greater institutional transparency enhances urban attractiveness to skilled labor. Such transparency supports human capital agglomeration and knowledge accumulation (Du, 2020). Overall, public data openness reshapes urban innovation by reconfiguring the institutional conditions under which capital, entrepreneurship, and talent are mobilized and coordinated.
Regarding theoretical contributions, this study conceptualizes public data openness as a distinct production factor and embeds it within an innovation framework. This study clarifies the mechanisms through which public data promotes urban innovation, including facilitating information flow, enhancing institutional transparency, and improving resource allocation efficiency. In doing so, the study extends the existing literature in a substantive way. By employing a staggered DID design and conducting a series of comprehensive robustness tests, the study provides credible causal evidence. This study systematically identifies the pathways through which public data openness affects urban innovation capability at the systemic level. These findings deepen the understanding of data as an institutionalized production factor. Furthermore, results address gaps in the literature on data openness policies, urban innovation, and their underlying mechanisms. Furthermore, the findings underscore the relevance of policy context and industry characteristics in shaping the transmission of these effects. Therefore, the study offers new perspectives and empirical support for research on data governance and urban innovation in the digital economy.
Policy implicationsFirst, to fully leverage the institutional role of public data openness in promoting urban innovation, policy implementation should prioritize enhancing data quality and strengthening governance. High-quality, standardized, and reliable public data is the foundation for activating investment, entrepreneurship, and talent mechanisms. Key actions include establishing unified data standards, improving cross-departmental data sharing, and implementing performance evaluation systems for open data platforms.
Second, the scope of pilot initiatives should be expanded strategically. In cities with advanced digital infrastructure and high economic openness, the breadth and depth of data openness should be increased. In less-developed regions, improvement of data collection, governance, and management systems should be emphasized for enhancing innovation capacity while supporting balanced regional development.
Third, public data openness should be coordinated with policies targeting investment, entrepreneurship, and human capital. By providing accessible, high-quality data, information acquisition costs are reduced, enabling more efficient investment and entrepreneurial decisions and fostering a virtuous cycle among these innovation drivers.
Finally, regions with underdeveloped infrastructure should strengthen their fiscal and technical support. Science and technology expenditure should allocate resources specifically for public data governance, platform operation, and application development. This promotes an in-depth integration of data with research and industry, amplifying the institutional and multiplier effects of public data on urban innovation, which can improve decision-making, enhance service delivery, and increase economic growth in urban areas.
DiscussionStudy limitationsThis study provides a systematic analysis of the impact of the public data openness pilot policy on urban innovation levels and explores the mediating mechanisms of the investment level, entrepreneurial activity, and human resource concentration. Nevertheless, several limitations should be acknowledged. First, the issue of selective openness in public data was not fully addressed. Governments often release data deemed secure and less sensitive, whereas genuinely high-value microlevel data remains largely closed to public access. Constrained by data availability, this study could not analyze the selectivity in data openness and its potential implications. Second, data granularity is relatively coarse, as most publicly available data is highly aggregated. This limits the precise measurement of innovation activities and may underestimate the policy’s true effect. Third, this study employs traditional metrics, such as patent counts, to measure urban innovation levels. While patent data primarily reflects formal technological innovation, it does not fully capture multidimensional innovative activities, including social innovation, service innovation, and digital innovation. Finally, the analysis is grounded in the single-country context of China’s pilot experience. The specific national background and the methodological scope may limit the external applicability of the findings. Variations in policy implementation pace, data openness levels, and market response mechanisms across different countries necessitate caution when extending this study’s conclusions to other institutional or regional contexts.
Future research directionsTo further advance research on the impact of open public data on urban innovation, future studies should explore three main directions. First, constructing a more refined evaluation index system is necessary for public data. This system should focus on the quantity of data and assess its criticality, accessibility, and update timeliness. By integrating text analysis and case studies, an indicator such as “data value density” could be established to improve the precision of measuring policy effects. Second, urban innovation outcomes can be analyzed across dimensions by distinguishing between incremental and disruptive innovation. Furthermore, research can investigate how data granularity influences different innovation mechanisms. Incorporating spatial econometric models to study knowledge spillovers and competitive effects among cities would be valuable. Finally, comparative cross-country studies should be conducted. Selecting countries with similar institutional or economic development levels but different data openness models, such as China and Singapore, may evaluate the applicability and generalizability of the policy across diverse institutional contexts. Such efforts would enrich the theoretical foundation and practical guidance on how public data openness influences urban innovation.
CRediT authorship contribution statementPinglu Zhou: Writing – review & editing, Visualization, Validation, Supervision, Software, Resources, Project administration, Methodology, Investigation, Funding acquisition, Formal analysis, Conceptualization. Wei Shi: Writing – review & editing, Writing – original draft, Visualization, Validation, Supervision, Software, Resources, Project administration, Methodology, Investigation, Funding acquisition, Formal analysis, Data curation, Conceptualization. Zhenyu Wang: Writing – review & editing, Visualization, Validation, Supervision, Software, Resources, Project administration, Methodology, Investigation, Funding acquisition, Formal analysis. Miao lv: Writing – review & editing.
The authors have no competing interests to declare that are relevant to the content of this article.
This research was supported by Soft Science Special Project of Gansu Basic Research Plan under Grant No. 24JRZA030.



























