• 文献检索
  • 文档翻译
  • 深度研究
  • 学术资讯
  • Suppr Zotero 插件Zotero 插件
  • 邀请有礼
  • 套餐&价格
  • 历史记录
应用&插件
Suppr Zotero 插件Zotero 插件浏览器插件Mac 客户端Windows 客户端微信小程序
定价
高级版会员购买积分包购买API积分包
服务
文献检索文档翻译深度研究API 文档MCP 服务
关于我们
关于 Suppr公司介绍联系我们用户协议隐私条款
关注我们

Suppr 超能文献

核心技术专利:CN118964589B侵权必究
粤ICP备2023148730 号-1Suppr @ 2026

文献检索

告别复杂PubMed语法,用中文像聊天一样搜索,搜遍4000万医学文献。AI智能推荐,让科研检索更轻松。

立即免费搜索

文件翻译

保留排版,准确专业,支持PDF/Word/PPT等文件格式,支持 12+语言互译。

免费翻译文档

深度研究

AI帮你快速写综述,25分钟生成高质量综述,智能提取关键信息,辅助科研写作。

立即免费体验

一项基于文本数据自动检测自闭症谱系障碍的跨数据集研究。

A cross-dataset study on automatic detection of autism spectrum disorder from text data.

作者信息

Wawer Aleksander, Chojnicka Izabela, Sarzyńska-Wawer Justyna, Krawczyk Małgorzata

机构信息

Polish Academy of Sciences, Institute of Computer Science, Warsaw, Poland.

Department of Health and Rehabilitation Psychology, Faculty of Psychology, University of Warsaw, Warsaw, Poland.

出版信息

Acta Psychiatr Scand. 2025 Mar;151(3):259-269. doi: 10.1111/acps.13737. Epub 2024 Jul 20.

DOI:10.1111/acps.13737
PMID:39032040
原文链接:https://pmc.ncbi.nlm.nih.gov/articles/PMC11787923/
Abstract

OBJECTIVE

The goals of this article are as follows. First, to investigate the possibility of detecting autism spectrum disorder (ASD) from text data using the latest generation of machine learning tools. Second, to compare model performance on two datasets of transcribed statements, collected using two different diagnostic tools. Third, to investigate the feasibility of knowledge transfer between models trained on both datasets and check if data augmentation can help alleviate the problem of a small number of observations.

METHOD

We explore two techniques to detect ASD. The first one is based on fine-tuning HerBERT, a BERT-based, monolingual deep transformer neural network. The second one uses the newest, multipurpose text embeddings from OpenAI and a classifier. We apply the methods to two separate datasets of transcribed statements, collected using two different diagnostic tools: thought, language, and communication (TLC) and autism diagnosis observation schedule-2 (ADOS-2). We conducted several cross-dataset experiments in both a zero-shot setting and a setting where models are pretrained on one dataset and then training continues on another to test the possibility of knowledge transfer.

RESULTS

Unlike previous studies, the models we tested obtained average results on ADOS-2 data but reached very good performance of the models in TLC. We did not observe any benefits from knowledge transfer between datasets. We observed relatively poor performance of models trained on augmented data and hypothesize that data augmentation by back translation obfuscates autism-specific signals.

CONCLUSION

The quality of machine learning models that detect ASD from text data is improving, but model results are dependent on the type of input data or diagnostic tool.

摘要

目的

本文的目标如下。其一,使用最新一代机器学习工具研究从文本数据中检测自闭症谱系障碍(ASD)的可能性。其二,比较在使用两种不同诊断工具收集的两个转录陈述数据集上的模型性能。其三,研究在两个数据集上训练的模型之间知识转移的可行性,并检查数据增强是否有助于缓解观测数据量少的问题。

方法

我们探索了两种检测ASD的技术。第一种基于对HerBERT(一种基于BERT的单语深度变换器神经网络)进行微调。第二种使用来自OpenAI的最新多用途文本嵌入和一个分类器。我们将这些方法应用于使用两种不同诊断工具收集的两个单独的转录陈述数据集:思维、语言和沟通(TLC)以及自闭症诊断观察量表-2(ADOS-2)。我们在零样本设置以及模型在一个数据集上进行预训练然后在另一个数据集上继续训练的设置下进行了几次跨数据集实验,以测试知识转移的可能性。

结果

与先前的研究不同,我们测试的模型在ADOS-2数据上获得了平均结果,但在TLC数据上模型达到了非常好的性能。我们没有观察到数据集之间知识转移带来的任何益处。我们观察到在增强数据上训练的模型性能相对较差,并推测通过反向翻译进行数据增强会模糊自闭症特异性信号。

结论

从文本数据中检测ASD的机器学习模型的质量正在提高,但模型结果取决于输入数据的类型或诊断工具。

https://cdn.ncbi.nlm.nih.gov/pmc/blobs/7d74/11787923/373e3d033659/ACPS-151-259-g001.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/7d74/11787923/373e3d033659/ACPS-151-259-g001.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/7d74/11787923/373e3d033659/ACPS-151-259-g001.jpg

相似文献

1
A cross-dataset study on automatic detection of autism spectrum disorder from text data.一项基于文本数据自动检测自闭症谱系障碍的跨数据集研究。
Acta Psychiatr Scand. 2025 Mar;151(3):259-269. doi: 10.1111/acps.13737. Epub 2024 Jul 20.
2
Memantine for autism spectrum disorder.美金刚治疗自闭症谱系障碍。
Cochrane Database Syst Rev. 2022 Aug 25;8(8):CD013845. doi: 10.1002/14651858.CD013845.pub2.
3
Signs and symptoms to determine if a patient presenting in primary care or hospital outpatient settings has COVID-19.在基层医疗机构或医院门诊环境中,如果患者出现以下症状和体征,可判断其是否患有 COVID-19。
Cochrane Database Syst Rev. 2022 May 20;5(5):CD013665. doi: 10.1002/14651858.CD013665.pub3.
4
Education support services for improving school engagement and academic performance of children and adolescents with a chronic health condition.改善患有慢性病的儿童和青少年的学校参与度和学业成绩的教育支持服务。
Cochrane Database Syst Rev. 2023 Feb 8;2(2):CD011538. doi: 10.1002/14651858.CD011538.pub2.
5
Leveraging a foundation model zoo for cell similarity search in oncological microscopy across devices.利用基础模型库进行跨设备肿瘤显微镜检查中的细胞相似性搜索。
Front Oncol. 2025 Jun 18;15:1480384. doi: 10.3389/fonc.2025.1480384. eCollection 2025.
6
Magnetic resonance perfusion for differentiating low-grade from high-grade gliomas at first presentation.首次就诊时磁共振灌注成像用于鉴别低级别与高级别胶质瘤
Cochrane Database Syst Rev. 2018 Jan 22;1(1):CD011551. doi: 10.1002/14651858.CD011551.pub2.
7
Individual-level interventions to reduce personal exposure to outdoor air pollution and their effects on people with long-term respiratory conditions.个体层面的干预措施以减少个人接触室外空气污染及其对长期呼吸系统疾病患者的影响。
Cochrane Database Syst Rev. 2021 Aug 9;8(8):CD013441. doi: 10.1002/14651858.CD013441.pub2.
8
Physical activity for treatment of irritable bowel syndrome.体力活动治疗肠易激综合征。
Cochrane Database Syst Rev. 2022 Jun 29;6(6):CD011497. doi: 10.1002/14651858.CD011497.pub2.
9
Falls prevention interventions for community-dwelling older adults: systematic review and meta-analysis of benefits, harms, and patient values and preferences.社区居住的老年人跌倒预防干预措施:系统评价和荟萃分析的益处、危害以及患者的价值观和偏好。
Syst Rev. 2024 Nov 26;13(1):289. doi: 10.1186/s13643-024-02681-3.
10
Gonadotropin-releasing hormone (GnRH) analogues for premenstrual syndrome (PMS).用于经前综合征(PMS)的促性腺激素释放激素(GnRH)类似物。
Cochrane Database Syst Rev. 2025 Jun 10;6(6):CD011330. doi: 10.1002/14651858.CD011330.pub2.

本文引用的文献

1
Autism spectrum disorder detection with kNN imputer and machine learning classifiers via questionnaire mode of screening.通过问卷调查筛选模式,利用k近邻插补法和机器学习分类器进行自闭症谱系障碍检测。
Health Inf Sci Syst. 2024 Mar 6;12(1):18. doi: 10.1007/s13755-024-00277-8. eCollection 2024 Dec.
2
Development of Deep Ensembles to Screen for Autism and Symptom Severity Using Retinal Photographs.利用视网膜照片开发深度集成模型以筛查自闭症及症状严重程度
JAMA Netw Open. 2023 Dec 1;6(12):e2347692. doi: 10.1001/jamanetworkopen.2023.47692.
3
A Review of and Roadmap for Data Science and Machine Learning for the Neuropsychiatric Phenotype of Autism.
自闭症神经精神表型的数据科学和机器学习综述及路线图。
Annu Rev Biomed Data Sci. 2023 Aug 10;6:211-228. doi: 10.1146/annurev-biodatasci-020722-125454. Epub 2023 May 3.
4
Prevalence and Characteristics of Autism Spectrum Disorder Among Children Aged 8 Years - Autism and Developmental Disabilities Monitoring Network, 11 Sites, United States, 2020.2020 年,美国 11 个监测点自闭症和发育障碍监测网络 8 岁儿童自闭症谱系障碍的流行率和特征。
MMWR Surveill Summ. 2023 Mar 24;72(2):1-14. doi: 10.15585/mmwr.ss7202a1.
5
Evaluation of a decided sample size in machine learning applications.机器学习应用中确定样本量的评估。
BMC Bioinformatics. 2023 Feb 14;24(1):48. doi: 10.1186/s12859-023-05156-9.
6
Machine learning based on eye-tracking data to identify Autism Spectrum Disorder: A systematic review and meta-analysis.基于眼动追踪数据的机器学习用于识别自闭症谱系障碍:一项系统综述和荟萃分析。
J Biomed Inform. 2023 Jan;137:104254. doi: 10.1016/j.jbi.2022.104254. Epub 2022 Dec 9.
7
Conventions for unconventional language: Revisiting a framework for spoken language features in autism.非传统语言的规范:重新审视自闭症口语特征的框架。
Autism Dev Lang Impair. 2022 Jun 5;7:23969415221105472. doi: 10.1177/23969415221105472. eCollection 2022 Jan-Dec.
8
Growth in narrative retelling and inference abilities and relations with reading comprehension in children and adolescents with autism spectrum disorder.自闭症谱系障碍儿童和青少年的叙事复述与推理能力发展及其与阅读理解的关系。
Autism Dev Lang Impair. 2020 Nov 20;5:2396941520968028. doi: 10.1177/2396941520968028. eCollection 2020 Jan-Dec.
9
Language and Speech Characteristics in Autism.自闭症中的语言和言语特征
Neuropsychiatr Dis Treat. 2022 Oct 14;18:2367-2377. doi: 10.2147/NDT.S331987. eCollection 2022.
10
Detecting autism from picture book narratives using deep neural utterance embeddings.使用深度神经网络话语嵌入来从绘本故事中检测自闭症。
Int J Lang Commun Disord. 2022 Sep;57(5):948-962. doi: 10.1111/1460-6984.12731. Epub 2022 May 12.