加载中...
Information Frictions, Public Data and Firms' Innovation Boundaries
在从“创新大国”迈向“创新强国”的战略转型中,如何破解突破式创新不足的结构性困境,已成为提升中国科技竞争力的关键命题。本文从信息摩擦视角出发,构建了一个包含私有数据与公共数据的异质性创新选择模型,刻画企业在渐进式与突破式创新之间的最优决策。理论分析表明,公共数据作为一种公共信息,一方面增强企业对外部创新环境的识别能力(信息供给),另一方面有助于降低企业间创新网络的协同成本(协调强化),从而提高企业研发效率并激励突破式创新,推动创新边界拓展。在此基础上,本文利用中国地方政府公共数据开放平台上线的冲击,对模型结论进行实证检验,结果与理论预测相符。进一步地,本文基于大语言模型对企业专利进行技术领域识别,计算其与地方政府公共数据所属领域的余弦相似度,构建技术布局引导变量,并结合企业未来研发投入波动率,共同验证了公共数据在技术引导和不确定性缓解方面的信息供给机制。此外,公共数据开放还显著推动了企业开放式创新,佐证了理论模型中的协调强化机制。本文拓展了信息经济学与创新经济学的交叉研究,为理解数据如何重构微观创新决策提供了理论基础与实证支撑,同时也为充分释放公共数据价值与完善政府数据开放平台建设提供了针对性的政策建议。
Abstract
In the digital era, public data has become a vital infrastructure for firm innovation. While governments globally are expanding public data sharing, how such data influences firms' strategic innovation, particularly under information frictions, remains poorly understood. In China, public data openness primarily involves low-frequency, unstructured information such as policy documents or industrial plans. Although these data are often unsuitable for direct algorithmic or production use, they provide substantial informational value by clarifying policy priorities, reducing uncertainty in the innovation landscape, and helping firms anticipate directional shifts in technological and regulatory environments. This paper investigates how public data influences the expansion of firms' innovation boundaries from the perspective of information frictions. Theoretically, we develop a model in which firms face uncertainty in R&D investment and must choose between incremental and radical innovation strategies. The model introduces a Bayesian signal environment where firms form expectations about the innovation landscape using both private and public data. Public data primarily operates through two channels: first, an information channel, whereby its informational content enhances firms' ability to anticipate innovation fundamentals and reduces uncertainty in R&D investment; and second, a coordination channel, whereby its shared nature allows multiple firms to utilize it collectively, mitigating inter-firm information asymmetries and improving collaborative efficiency. Empirically, we exploit the staggered rollout of municipal public data platforms in China as a quasi-natural experiment. The results show that the launch of public data platforms significantly promotes firms' radical innovation. Mechanism analyses further reveal that public data access reduces R&D investment volatility and guides firms' technological deployment toward policy-relevant domains, supporting the information channel. In addition, public data facilitates inter-organizational collaboration by enhancing coordination, as evidenced by increased joint patenting, thereby validating the coordination channel. This study makes three main contributions. First, it develops a theoretical model of heterogeneous innovation choices under information frictions, integrating both private and public data into firms' R&D decision-making. Second, this paper offers a new perspective on the determinants of radical innovation. Existing studies have focused on factors such as managerial characteristics, government subsidies, and technological spillovers. In contrast, this study emphasizes the information environment as a key driver of radical innovation by identifying the information and coordination effects of public data. Third, it contributes to the growing literature on public data policy evaluation. This paper employs large language models to extract and classify unstructured content from both patent texts and public data directories. This enables a novel measure of topic-level alignment between government-supplied data and firms' innovation activities, offering new evidence on the micro-level channels through which public data affects firm behavior. This paper advances the interdisciplinary research at the intersection of information economics and innovation boundary theory, providing a solid theoretical foundation and empirical evidence for understanding how data, by alleviating information frictions, reconfigures firms' micro-level innovation decisions. Furthermore, the findings offer targeted policy insights for fully unlocking the informational value of public data and optimizing the functional design of government data openness platforms. The study also offers novel perspectives and practical implications for leveraging data as a key factor to empower high-quality economic development, modernize national governance, and advance digital government development.
公共数据开放显著提升企业创新边界,突破式创新比例平均提高1.7个百分点
以地方政府公共数据开放平台上线的双重差分估计,企业创新边界(突破式创新专利占比)在平台上线后显著提高1.7个百分点,相当于样本均值的16%,表明公共数据开放有效推动了企业将研发资源配置向突破式创新倾斜。
公共数据开放显著降低企业研发投资波动率,缓解创新不确定性
以企业研发投入占营业收入比例的三期滚动标准差度量R&D投资波动率,发现公共数据开放显著降低了企业R&D投资波动率,且剔除了政府补助影响后的净R&D波动率同样显著下降,表明信息供给机制有效缓解了创新活动中的不确定性。
公共数据开放引导企业技术布局向政府数据开放的重点领域倾斜
基于大语言模型将企业专利文本与城市公共数据条目进行主题匹配,构造技术布局引导变量,发现公共数据开放显著提高了企业未来专利与政府开放数据领域的一致性,说明公共数据发挥了技术引导的信号功能。
公共数据开放显著提升合作专利产出,促进开放式创新与产学研协同
以联合申请专利衡量开放式创新,发现公共数据开放显著提升了企业合作专利数量和产学研合作专利产出,而非合作专利不显著,验证了协调强化机制:共享的公共数据降低了创新网络中的沟通与信任成本,促进了协同创新。
核心解释变量
公共数据开放(地方政府数据开放平台上线的政策虚拟变量DID)
被解释变量
企业创新边界(突破式创新专利占同期所有申请专利的比例)
样本与数据
2007-2022年中国沪深A股上市公司数据,剔除金融行业、资产负债率大于1、当年ST/PT、样本期间企业办公地点变化及主要变量缺失的样本,连续变量进行上下1%缩尾处理,标准误聚类到城市层面
识别方法 / 模型设定
以地方政府公共数据开放平台上线的准自然实验,构建双重差分模型(DID),控制企业固定效应和行业-年份固定效应,并控制企业层面和城市层面特征变量
内生性及稳健性检验
进行了多维度的稳健性与异质性检验,包括:动态效应与平行趋势检验、排除竞争性假设(公共数据作为生产要素)、溢出效应讨论(限定小市场城市、剔除直辖市/副省级城市、剔除处理组相邻城市、异地经营规模较小企业子样本)、更换变量度量方式、替换公共数据开放度量方式(普及程度)、引入先验信息和接收者噪声的扩展模型分析等,核心结论稳健成立
A股上市公司专利及引用被引用数据
上市公司可支持的研究问题:可用于获取A股上市公司专利的引用与被引用数据,深入分析公共数据开放后企业技术知识流动和创新质量的变化。
中国全部专利申请与授权数据
专利申请可支持的研究问题:该数据集提供中国全部专利申请与授权信息,可用于识别和计量企业突破式专利(新技术领域)的数量与结构特征。
中国绿色低碳专利申请与授权数据
绿色低碳可支持的研究问题:该数据集可用于研究地方政府开放特定领域(如绿色低碳技术)的数据后,是否显著引导和激励了企业在该领域的创新活动。
中国各地区政府工作报告文本数据
政府工作报告可支持的研究问题:该数据库提供各级政府工作报告全文,可用于提取政策导向和重点发展领域信息,并将之与企业创新方向相结合进行分析。
A股上市公司及子公司基本信息扩展数据
上市公司可支持的研究问题:可用于识别企业异地子公司分布,评估数据开放的跨区域溢出效应,以及企业集团内部的创新网络特征。