The Alan Turing Institute, London, United Kingdom.
University of Amsterdam, Amsterdam, Netherlands.
PLoS One. 2020 Apr 22;15(4):e0230416. doi: 10.1371/journal.pone.0230416. eCollection 2020.
Efforts to make research results open and reproducible are increasingly reflected by journal policies encouraging or mandating authors to provide data availability statements. As a consequence of this, there has been a strong uptake of data availability statements in recent literature. Nevertheless, it is still unclear what proportion of these statements actually contain well-formed links to data, for example via a URL or permanent identifier, and if there is an added value in providing such links. We consider 531, 889 journal articles published by PLOS and BMC, develop an automatic system for labelling their data availability statements according to four categories based on their content and the type of data availability they display, and finally analyze the citation advantage of different statement categories via regression. We find that, following mandated publisher policies, data availability statements become very common. In 2018 93.7% of 21,793 PLOS articles and 88.2% of 31,956 BMC articles had data availability statements. Data availability statements containing a link to data in a repository-rather than being available on request or included as supporting information files-are a fraction of the total. In 2017 and 2018, 20.8% of PLOS publications and 12.2% of BMC publications provided DAS containing a link to data in a repository. We also find an association between articles that include statements that link to data in a repository and up to 25.36% (± 1.07%) higher citation impact on average, using a citation prediction model. We discuss the potential implications of these results for authors (researchers) and journal publishers who make the effort of sharing their data in repositories. All our data and code are made available in order to reproduce and extend our results.
为了使研究结果开放和可重现,期刊政策越来越鼓励或要求作者提供数据可用性声明。因此,最近的文献中大量采用了数据可用性声明。然而,目前尚不清楚这些声明中有多少实际上包含了指向数据的形式良好的链接,例如通过 URL 或永久标识符,以及提供此类链接是否有额外的价值。我们考虑了 PLOS 和 BMC 出版的 531,889 篇期刊文章,根据其内容和显示的数据可用性类型,开发了一种自动系统,根据四个类别对其数据可用性声明进行标记,并最终通过回归分析不同声明类别的引用优势。我们发现,根据强制出版商政策,数据可用性声明变得非常普遍。2018 年,PLOS 的 21,793 篇文章中有 93.7%,BMC 的 31,956 篇文章中有 88.2%有数据可用性声明。包含指向存储库中数据的链接的数据可用性声明而不是按需提供或包含在支持信息文件中的声明仅占总数的一小部分。2017 年和 2018 年,PLOS 出版物中有 20.8%,BMC 出版物中有 12.2%提供了包含指向存储库中数据的链接的 DAS。我们还发现,包含指向存储库中数据的链接的声明与平均高达 25.36%(±1.07%)的引用影响力之间存在关联,使用引文预测模型。我们讨论了这些结果对作者(研究人员)和发布者共享数据的潜在影响。我们提供了所有的数据和代码,以便复制和扩展我们的结果。