Suppr超能文献

性别检测工具的性能:姓名到性别推断服务的比较研究。

Performance of gender detection tools: a comparative study of name-to-gender inference services.

出版信息

J Med Libr Assoc. 2021 Jul 1;109(3):414-421. doi: 10.5195/jmla.2021.1185.

Abstract

OBJECTIVE

To evaluate the performance of gender detection tools that allow the uploading of files (e.g., Excel or CSV files) containing first names, are usable by researchers without advanced computer skills, and are at least partially free of charge.

METHODS

The study was conducted using four physician datasets (total number of physicians: 6,131; 50.3% female) from Switzerland, a multilingual country. Four gender detection tools met the inclusion criteria: three partially free (Gender API, NamSor, and genderize.io) and one completely free (Wiki-Gendersort). For each tool, we recorded the number of correct classifications (i.e., correct gender assigned to a name), misclassifications (i.e., wrong gender assigned to a name), and nonclassifications (i.e., no gender assigned). We computed three metrics: the proportion of misclassifications excluding nonclassifications (errorCodedWithoutNA), the proportion of nonclassifications (naCoded), and the proportion of misclassifications and nonclassifications (errorCoded).

RESULTS

The proportion of misclassifications was low for all four gender detection tools (errorCodedWithoutNA between 1.5 and 2.2%). By contrast, the proportion of unrecognized names (naCoded) varied: 0% for NamSor, 0.3% for Gender API, 4.5% for Wiki-Gendersort, and 16.4% for genderize.io. Using errorCoded, which penalizes both types of error equally, we obtained the following results: Gender API 1.8%, NamSor 2.0%, Wiki-Gendersort 6.6%, and genderize.io 17.7%.

CONCLUSIONS

Gender API and NamSor were the most accurate tools. Genderize.io led to a high number of nonclassifications. Wiki-Gendersort may be a good compromise for researchers wishing to use a completely free tool. Other studies would be useful to evaluate the performance of these tools in other populations (e.g., Asian).

摘要

目的

评估允许上传包含名字的文件(如 Excel 或 CSV 文件)的性别检测工具的性能,这些工具可供没有高级计算机技能的研究人员使用,且至少部分免费。

方法

本研究使用了来自瑞士(一个多语言国家)的四个医生数据集(医生总数:6131 人;女性占 50.3%)。有四个性别检测工具符合纳入标准:三个部分免费(Gender API、NamSor 和 genderize.io)和一个完全免费(Wiki-Gendersort)。对于每个工具,我们记录了正确分类的数量(即正确分配给一个名字的性别)、错误分类的数量(即错误分配给一个名字的性别)和未分类的数量。我们计算了三个指标:排除未分类的错误分类比例(errorCodedWithoutNA)、未分类的比例(naCoded)和错误分类和未分类的比例(errorCoded)。

结果

所有四个性别检测工具的错误分类比例都较低(errorCodedWithoutNA 在 1.5%到 2.2%之间)。相比之下,未识别名字的比例(naCoded)有所不同:NamSor 为 0%,Gender API 为 0.3%,Wiki-Gendersort 为 4.5%,genderize.io 为 16.4%。使用同样惩罚两种错误的 errorCoded,我们得到以下结果:Gender API 为 1.8%,NamSor 为 2.0%,Wiki-Gendersort 为 6.6%,genderize.io 为 17.7%。

结论

Gender API 和 NamSor 是最准确的工具。genderize.io 导致了大量的未分类。对于希望使用完全免费工具的研究人员来说,Wiki-Gendersort 可能是一个很好的折中方案。其他研究将有助于评估这些工具在其他人群(如亚洲)中的性能。

相似文献

7
Comparison and benchmark of name-to-gender inference services.姓名到性别的推理服务的比较与基准测试
PeerJ Comput Sci. 2018 Jul 16;4:e156. doi: 10.7717/peerj-cs.156. eCollection 2018.

引用本文的文献

本文引用的文献

1
Comparison and benchmark of name-to-gender inference services.姓名到性别的推理服务的比较与基准测试
PeerJ Comput Sci. 2018 Jul 16;4:e156. doi: 10.7717/peerj-cs.156. eCollection 2018.
2
Gender disparities in coronavirus disease 2019 clinical trial leadership.2019 年冠状病毒病临床试验领导中的性别差异。
Clin Microbiol Infect. 2021 Jul;27(7):1007-1010. doi: 10.1016/j.cmi.2020.12.025. Epub 2021 Jan 5.
4
Sex Distribution of Editorial Board Members Among Emergency Medicine Journals.急诊医学期刊编辑委员会成员的性别分布。
Ann Emerg Med. 2021 Jan;77(1):117-123. doi: 10.1016/j.annemergmed.2020.03.027. Epub 2020 May 4.
6
Sex and gender reporting in global health: new editorial policies.全球健康领域中的性别与性取向报告:新编辑政策
BMJ Glob Health. 2018 Jul 26;3(4):e001038. doi: 10.1136/bmjgh-2018-001038. eCollection 2018.
8
Gendermetrics of cancer research: results from a global analysis on lung cancer.癌症研究的性别指标:肺癌全球分析结果
Oncotarget. 2017 Oct 26;8(60):101911-101921. doi: 10.18632/oncotarget.22089. eCollection 2017 Nov 24.

文献检索

告别复杂PubMed语法,用中文像聊天一样搜索,搜遍4000万医学文献。AI智能推荐,让科研检索更轻松。

立即免费搜索

文件翻译

保留排版,准确专业,支持PDF/Word/PPT等文件格式,支持 12+语言互译。

免费翻译文档

深度研究

AI帮你快速写综述,25分钟生成高质量综述,智能提取关键信息,辅助科研写作。

立即免费体验