• 文献检索
  • 文档翻译
  • 深度研究
  • 学术资讯
  • Suppr Zotero 插件Zotero 插件
  • 邀请有礼
  • 套餐&价格
  • 历史记录
应用&插件
Suppr Zotero 插件Zotero 插件浏览器插件Mac 客户端Windows 客户端微信小程序
定价
高级版会员购买积分包购买API积分包
服务
文献检索文档翻译深度研究API 文档MCP 服务
关于我们
关于 Suppr公司介绍联系我们用户协议隐私条款
关注我们

Suppr 超能文献

核心技术专利:CN118964589B侵权必究
粤ICP备2023148730 号-1Suppr @ 2026

文献检索

告别复杂PubMed语法,用中文像聊天一样搜索,搜遍4000万医学文献。AI智能推荐,让科研检索更轻松。

立即免费搜索

文件翻译

保留排版,准确专业,支持PDF/Word/PPT等文件格式,支持 12+语言互译。

免费翻译文档

深度研究

AI帮你快速写综述,25分钟生成高质量综述,智能提取关键信息,辅助科研写作。

立即免费体验

减轻CORD-19中用于分析COVID-19文献的偏差。

Mitigating Biases in CORD-19 for Analyzing COVID-19 Literature.

作者信息

Kanakia Anshul, Wang Kuansan, Dong Yuxiao, Xie Boya, Lo Kyle, Shen Zhihong, Wang Lucy Lu, Huang Chiyuan, Eide Darrin, Kohlmeier Sebastian, Wu Chieh-Han

机构信息

Microsoft Research, Redmond, WA, United States.

Allen Institute for Artificial Intelligence, Seattle, WA, United States.

出版信息

Front Res Metr Anal. 2020 Nov 23;5:596624. doi: 10.3389/frma.2020.596624. eCollection 2020.

DOI:10.3389/frma.2020.596624
PMID:33870059
原文链接:https://pmc.ncbi.nlm.nih.gov/articles/PMC8025972/
Abstract

On the behest of the Office of Science and Technology Policy in the White House, six institutions, including ours, have created an open research dataset called COVID-19 Research Dataset (CORD-19) to facilitate the development of question-answering systems that can assist researchers in finding relevant research on COVID-19. As of May 27, 2020, CORD-19 includes more than 100,000 open access publications from major publishers and PubMed as well as preprint articles deposited into medRxiv, bioRxiv, and arXiv. Recent years, however, have also seen question-answering and other machine learning systems exhibit harmful behaviors to humans due to biases in the training data. It is imperative and only ethical for modern scientists to be vigilant in inspecting and be prepared to mitigate the potential biases when working with any datasets. This article describes a framework to examine biases in scientific document collections like CORD-19 by comparing their properties with those derived from the citation behaviors of the entire scientific community. In total, three expanded sets are created for the analyses: 1) the enclosure set CORD-19E composed of CORD-19 articles and their references and citations, mirroring the methodology used in the renowned "A Century of Physics" analysis; 2) the full closure graph CORD-19C that recursively includes references starting with CORD-19; and 3) the inflection closure CORD-19I, that is, a much smaller subset of CORD-19C but already appropriate for statistical analysis based on the theory of the scale-free nature of the citation network. Taken together, all these expanded datasets show much smoother trends when used to analyze global COVID-19 research. The results suggest that while CORD-19 exhibits a strong tilt toward recent and topically focused articles, the knowledge being explored to attack the pandemic encompasses a much longer time span and is very interdisciplinary. A question-answering system with such expanded scope of knowledge may perform better in understanding the literature and answering related questions. However, while CORD-19 appears to have topical coverage biases compared to the expanded sets, the collaboration patterns, especially in terms of team sizes and geographical distributions, are captured very well already in CORD-19 as the raw statistics and trends agree with those from larger datasets.

摘要

应白宫科学技术政策办公室的要求,包括我们机构在内的六个机构创建了一个名为“COVID-19研究数据集(CORD-19)”的开放研究数据集,以促进问答系统的开发,该系统可以帮助研究人员查找有关COVID-19的相关研究。截至2020年5月27日,CORD-19包含来自主要出版商和PubMed的超过100,000篇开放获取出版物,以及存入medRxiv、bioRxiv和arXiv的预印本文章。然而,近年来,由于训练数据中的偏差,问答系统和其他机器学习系统也出现了对人类有害的行为。对于现代科学家来说,在使用任何数据集时保持警惕,检查并准备减轻潜在偏差是必要且符合道德规范的。本文描述了一个框架,通过将科学文献集合(如CORD-19)的属性与从整个科学界的引用行为中得出的属性进行比较,来检查其中的偏差。总共创建了三个扩展集用于分析:1)封闭集CORD-19E,由CORD-19文章及其参考文献和引用组成,反映了著名的“物理学百年”分析中使用的方法;2)完全封闭图CORD-19C,它递归地包含以CORD-19开头的参考文献;3)拐点封闭集CORD-19I,即CORD-19C的一个小得多的子集,但已适合基于引用网络的无标度性质理论进行统计分析。综合来看,所有这些扩展数据集在用于分析全球COVID-19研究时显示出更平滑的趋势。结果表明,虽然CORD-19对近期和主题聚焦的文章有强烈倾向,但用于应对大流行所探索的知识涵盖了更长的时间跨度且非常跨学科。一个具有如此扩展知识范围的问答系统在理解文献和回答相关问题方面可能表现得更好。然而,虽然与扩展集相比,CORD-19似乎存在主题覆盖偏差,但协作模式,特别是在团队规模和地理分布方面,在CORD-19中已经很好地体现出来了,因为原始统计数据和趋势与来自更大数据集的数据一致。

https://cdn.ncbi.nlm.nih.gov/pmc/blobs/1fc6/8025972/4c94388fefee/frma-05-596624-g014.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/1fc6/8025972/6e57fb47a64d/frma-05-596624-g001.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/1fc6/8025972/5853fd4d6eda/frma-05-596624-g002.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/1fc6/8025972/61a1e1bbd60c/frma-05-596624-g003.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/1fc6/8025972/28e1f6e1a1ef/frma-05-596624-g004.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/1fc6/8025972/6a2b45f9ccde/frma-05-596624-g005.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/1fc6/8025972/342105fb127d/frma-05-596624-g006.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/1fc6/8025972/62787759d5c0/frma-05-596624-g007.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/1fc6/8025972/27ca963677a9/frma-05-596624-g008.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/1fc6/8025972/5ba05791b5ff/frma-05-596624-g009.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/1fc6/8025972/bdc605aa16a6/frma-05-596624-g010.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/1fc6/8025972/6aeac6f275d6/frma-05-596624-g011.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/1fc6/8025972/244e30d3a5b5/frma-05-596624-g012.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/1fc6/8025972/f200fcc69759/frma-05-596624-g013.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/1fc6/8025972/4c94388fefee/frma-05-596624-g014.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/1fc6/8025972/6e57fb47a64d/frma-05-596624-g001.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/1fc6/8025972/5853fd4d6eda/frma-05-596624-g002.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/1fc6/8025972/61a1e1bbd60c/frma-05-596624-g003.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/1fc6/8025972/28e1f6e1a1ef/frma-05-596624-g004.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/1fc6/8025972/6a2b45f9ccde/frma-05-596624-g005.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/1fc6/8025972/342105fb127d/frma-05-596624-g006.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/1fc6/8025972/62787759d5c0/frma-05-596624-g007.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/1fc6/8025972/27ca963677a9/frma-05-596624-g008.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/1fc6/8025972/5ba05791b5ff/frma-05-596624-g009.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/1fc6/8025972/bdc605aa16a6/frma-05-596624-g010.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/1fc6/8025972/6aeac6f275d6/frma-05-596624-g011.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/1fc6/8025972/244e30d3a5b5/frma-05-596624-g012.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/1fc6/8025972/f200fcc69759/frma-05-596624-g013.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/1fc6/8025972/4c94388fefee/frma-05-596624-g014.jpg

相似文献

1
Mitigating Biases in CORD-19 for Analyzing COVID-19 Literature.减轻CORD-19中用于分析COVID-19文献的偏差。
Front Res Metr Anal. 2020 Nov 23;5:596624. doi: 10.3389/frma.2020.596624. eCollection 2020.
2
Folic acid supplementation and malaria susceptibility and severity among people taking antifolate antimalarial drugs in endemic areas.在流行地区,服用抗叶酸抗疟药物的人群中,叶酸补充剂与疟疾易感性和严重程度的关系。
Cochrane Database Syst Rev. 2022 Feb 1;2(2022):CD014217. doi: 10.1002/14651858.CD014217.
3
The future of Cochrane Neonatal.考克兰新生儿协作网的未来。
Early Hum Dev. 2020 Nov;150:105191. doi: 10.1016/j.earlhumdev.2020.105191. Epub 2020 Sep 12.
4
Social Media and Research Publication Activity During Early Stages of the COVID-19 Pandemic: Longitudinal Trend Analysis.社交媒体与 COVID-19 大流行早期阶段的研究出版活动:纵向趋势分析。
J Med Internet Res. 2021 Jun 17;23(6):e26956. doi: 10.2196/26956.
5
The Citation Cloud of a biomedical article: a free, public, web-based tool enabling citation analysis.生物医学文章的引文云:一个免费的、公共的、基于网络的工具,实现引文分析。
J Med Libr Assoc. 2022 Jan 1;110(1):103-108. doi: 10.5195/jmla.2022.1117.
6
Macromolecular crowding: chemistry and physics meet biology (Ascona, Switzerland, 10-14 June 2012).大分子拥挤现象:化学与物理邂逅生物学(瑞士阿斯科纳,2012年6月10日至14日)
Phys Biol. 2013 Aug;10(4):040301. doi: 10.1088/1478-3975/10/4/040301. Epub 2013 Aug 2.
7
A Comprehensive Overview of the COVID-19 Literature: Machine Learning-Based Bibliometric Analysis.《COVID-19 文献综述:基于机器学习的文献计量分析》
J Med Internet Res. 2021 Mar 8;23(3):e23703. doi: 10.2196/23703.
8
Research Trends in the Application of Artificial Intelligence in Oncology: A Bibliometric and Network Visualization Study.人工智能在肿瘤学应用中的研究趋势:文献计量学和网络可视化研究。
Front Biosci (Landmark Ed). 2022 Aug 31;27(9):254. doi: 10.31083/j.fbl2709254.
9
Scientific basis of the OCRA method for risk assessment of biomechanical overload of upper limb, as preferred method in ISO standards on biomechanical risk factors.OCRA 方法评估上肢生物力学过载风险的科学基础,作为 ISO 生物力学风险因素标准中的首选方法。
Scand J Work Environ Health. 2018 Jul 1;44(4):436-438. doi: 10.5271/sjweh.3746.
10
Discovering temporal scientometric knowledge in COVID-19 scholarly production.在新冠疫情学术成果中发现时间性科学计量学知识。
Scientometrics. 2022;127(3):1609-1642. doi: 10.1007/s11192-021-04260-y. Epub 2022 Jan 16.

引用本文的文献

1
Understanding progress in software citation: a study of software citation in the CORD-19 corpus.理解软件引用方面的进展:对CORD-19语料库中软件引用的研究。
PeerJ Comput Sci. 2022 Jul 25;8:e1022. doi: 10.7717/peerj-cs.1022. eCollection 2022.
2
Visibility, collaboration and impact of the Cuban scientific output on COVID-19 in Scopus.古巴科学成果在Scopus中关于新冠疫情研究的可见性、合作性及影响力
Heliyon. 2021 Oct 27;7(11):e08258. doi: 10.1016/j.heliyon.2021.e08258. eCollection 2021 Nov.
3
The boundary-spanning mechanisms of Nobel Prize winning papers.

本文引用的文献

1
Artificial-intelligence tools aim to tame the coronavirus literature.人工智能工具旨在梳理新冠病毒相关文献。
Nature. 2020 Jun 9. doi: 10.1038/d41586-020-01733-7.
2
A Review of Microsoft Academic Services for Science of Science Studies.微软学术服务在科学学研究方面的综述。
Front Big Data. 2019 Dec 3;2:45. doi: 10.3389/fdata.2019.00045. eCollection 2019.
3
A scientometric overview of CORD-19.CORD-19 的科学计量学概述。
诺贝尔奖获奖论文的跨界机制。
PLoS One. 2021 Aug 11;16(8):e0254744. doi: 10.1371/journal.pone.0254744. eCollection 2021.
4
A scientometric overview of CORD-19.CORD-19 的科学计量学概述。
PLoS One. 2021 Jan 7;16(1):e0244839. doi: 10.1371/journal.pone.0244839. eCollection 2021.
PLoS One. 2021 Jan 7;16(1):e0244839. doi: 10.1371/journal.pone.0244839. eCollection 2021.
4
Rare and everywhere: Perspectives on scale-free networks.稀有且无处不在:无标度网络的视角。
Nat Commun. 2019 Mar 4;10(1):1016. doi: 10.1038/s41467-019-09038-8.
5
Scale-free networks are rare.无标度网络很罕见。
Nat Commun. 2019 Mar 4;10(1):1017. doi: 10.1038/s41467-019-08746-5.
6
Maps of random walks on complex networks reveal community structure.复杂网络上随机游走的图谱揭示了群落结构。
Proc Natl Acad Sci U S A. 2008 Jan 29;105(4):1118-23. doi: 10.1073/pnas.0706851105. Epub 2008 Jan 23.
7
NETWORKS OF SCIENTIFIC PAPERS.科学论文网络
Science. 1965 Jul 30;149(3683):510-5. doi: 10.1126/science.149.3683.510.
8
Scale-free networks from varying vertex intrinsic fitness.来自不同顶点内在适应性的无标度网络。
Phys Rev Lett. 2002 Dec 16;89(25):258702. doi: 10.1103/PhysRevLett.89.258702. Epub 2002 Dec 3.
9
Community structure in social and biological networks.社会和生物网络中的群落结构。
Proc Natl Acad Sci U S A. 2002 Jun 11;99(12):7821-6. doi: 10.1073/pnas.122653799.
10
Emergence of scaling in random networks.随机网络中幂律分布的出现。
Science. 1999 Oct 15;286(5439):509-12. doi: 10.1126/science.286.5439.509.