• 文献检索
  • 文档翻译
  • 深度研究
  • 学术资讯
  • Suppr Zotero 插件Zotero 插件
  • 邀请有礼
  • 套餐&价格
  • 历史记录
应用&插件
Suppr Zotero 插件Zotero 插件浏览器插件Mac 客户端Windows 客户端微信小程序
定价
高级版会员购买积分包购买API积分包
服务
文献检索文档翻译深度研究API 文档MCP 服务
关于我们
关于 Suppr公司介绍联系我们用户协议隐私条款
关注我们

Suppr 超能文献

核心技术专利:CN118964589B侵权必究
粤ICP备2023148730 号-1Suppr @ 2026

文献检索

告别复杂PubMed语法,用中文像聊天一样搜索,搜遍4000万医学文献。AI智能推荐,让科研检索更轻松。

立即免费搜索

文件翻译

保留排版,准确专业,支持PDF/Word/PPT等文件格式,支持 12+语言互译。

免费翻译文档

深度研究

AI帮你快速写综述,25分钟生成高质量综述,智能提取关键信息,辅助科研写作。

立即免费体验

主动倾听。

Active listening.

机构信息

The Wellcome Centre for Human Neuroimaging, UCL Queen Square Institute of Neurology, London, WC1N 3AR, UK.

出版信息

Hear Res. 2021 Jan;399:107998. doi: 10.1016/j.heares.2020.107998. Epub 2020 May 20.

DOI:10.1016/j.heares.2020.107998
PMID:32732017
原文链接:https://pmc.ncbi.nlm.nih.gov/articles/PMC7812378/
Abstract

This paper introduces active listening, as a unified framework for synthesising and recognising speech. The notion of active listening inherits from active inference, which considers perception and action under one universal imperative: to maximise the evidence for our (generative) models of the world. First, we describe a generative model of spoken words that simulates (i) how discrete lexical, prosodic, and speaker attributes give rise to continuous acoustic signals; and conversely (ii) how continuous acoustic signals are recognised as words. The 'active' aspect involves (covertly) segmenting spoken sentences and borrows ideas from active vision. It casts speech segmentation as the selection of internal actions, corresponding to the placement of word boundaries. Practically, word boundaries are selected that maximise the evidence for an internal model of how individual words are generated. We establish face validity by simulating speech recognition and showing how the inferred content of a sentence depends on prior beliefs and background noise. Finally, we consider predictive validity by associating neuronal or physiological responses, such as the mismatch negativity and P300, with belief updating under active listening, which is greatest in the absence of accurate prior beliefs about what will be heard next.

摘要

本文介绍了主动倾听,作为一种将语音综合和识别的统一框架。主动倾听的概念源自主动推断,它将感知和行动置于一个普遍的准则之下:最大限度地提高我们对世界生成模型的证据。首先,我们描述了一个口语单词的生成模型,该模型模拟了(i)离散的词汇、韵律和说话人属性如何产生连续的声学信号;以及相反地(ii)如何将连续的声学信号识别为单词。“主动”方面涉及(隐蔽地)分割口语句子,并借鉴主动视觉的思想。它将语音分割看作是内部动作的选择,对应于单词边界的位置。实际上,选择的单词边界可以最大限度地提高关于单词生成方式的内部模型的证据。我们通过模拟语音识别来建立表面有效性,并展示句子的推断内容如何取决于先验信念和背景噪声。最后,我们通过将神经元或生理反应(如失匹配负波和 P300)与主动倾听下的信念更新相关联来考虑预测有效性,在缺乏对接下来会听到的内容的准确先验信念的情况下,这种更新最为强烈。

https://cdn.ncbi.nlm.nih.gov/pmc/blobs/1f3e/7812378/8cd65f423a2b/fx2.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/1f3e/7812378/fa550d63d904/gr1.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/1f3e/7812378/47e42fd37dac/gr2.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/1f3e/7812378/a62b4732c5d2/gr3.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/1f3e/7812378/e992b23ca39b/gr4.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/1f3e/7812378/bc1615dc4303/gr5.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/1f3e/7812378/0f1d914558ab/gr6.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/1f3e/7812378/d89a57bb874d/gr7.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/1f3e/7812378/3ca7e2090c22/gr8.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/1f3e/7812378/45b2468f2c14/gr9.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/1f3e/7812378/89a0eece32b0/gr10.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/1f3e/7812378/ee8aade1ab9f/fx1.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/1f3e/7812378/8cd65f423a2b/fx2.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/1f3e/7812378/fa550d63d904/gr1.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/1f3e/7812378/47e42fd37dac/gr2.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/1f3e/7812378/a62b4732c5d2/gr3.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/1f3e/7812378/e992b23ca39b/gr4.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/1f3e/7812378/bc1615dc4303/gr5.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/1f3e/7812378/0f1d914558ab/gr6.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/1f3e/7812378/d89a57bb874d/gr7.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/1f3e/7812378/3ca7e2090c22/gr8.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/1f3e/7812378/45b2468f2c14/gr9.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/1f3e/7812378/89a0eece32b0/gr10.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/1f3e/7812378/ee8aade1ab9f/fx1.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/1f3e/7812378/8cd65f423a2b/fx2.jpg

相似文献

1
Active listening.主动倾听。
Hear Res. 2021 Jan;399:107998. doi: 10.1016/j.heares.2020.107998. Epub 2020 May 20.
2
Extrinsic Cognitive Load Impairs Spoken Word Recognition in High- and Low-Predictability Sentences.外在认知负荷会影响高低预测度句子中的口语词汇识别。
Ear Hear. 2018 Mar/Apr;39(2):378-389. doi: 10.1097/AUD.0000000000000493.
3
Some Neurocognitive Correlates of Noise-Vocoded Speech Perception in Children With Normal Hearing: A Replication and Extension of ).听力正常儿童噪声-声码语音感知的一些神经认知关联:一项(研究的)复制与扩展 。 (注:原文括号部分不完整,翻译时保留原样)
Ear Hear. 2017 May/Jun;38(3):344-356. doi: 10.1097/AUD.0000000000000393.
4
Tracking Cognitive Spare Capacity During Speech Perception With EEG/ERP: Effects of Cognitive Load and Sentence Predictability.使用 EEG/ERP 追踪言语感知过程中的认知备用容量:认知负荷和句子可预测性的影响。
Ear Hear. 2020 Sep/Oct;41(5):1144-1157. doi: 10.1097/AUD.0000000000000856.
5
Effect of training on word-recognition performance in noise for young normal-hearing and older hearing-impaired listeners.训练对年轻听力正常者和老年听力受损者在噪声环境下单词识别能力的影响。
Ear Hear. 2006 Jun;27(3):263-78. doi: 10.1097/01.aud.0000215980.21158.a2.
6
Syllable Inference as a Mechanism for Spoken Language Understanding.音节推断作为口语理解的一种机制。
Top Cogn Sci. 2021 Apr;13(2):351-398. doi: 10.1111/tops.12529. Epub 2021 Mar 29.
7
Speech Perception in Noise and Listening Effort of Older Adults With Nonlinear Frequency Compression Hearing Aids.老年人使用非线性频率压缩助听器的噪声言语感知和聆听努力。
Ear Hear. 2018 Mar/Apr;39(2):215-225. doi: 10.1097/AUD.0000000000000481.
8
Errors on a Speech-in-Babble Sentence Recognition Test Reveal Individual Differences in Acoustic Phonetic Perception and Babble Misallocations.嘈杂语音句子识别测试中的错误揭示了声学语音感知和嘈杂语音误分配方面的个体差异。
Ear Hear. 2021 May/Jun;42(3):673-690. doi: 10.1097/AUD.0000000000001020.
9
Predictive Sentence Context Reduces Listening Effort in Older Adults With and Without Hearing Loss and With High and Low Working Memory Capacity.预测句子语境可减少有听力损失和无听力损失老年人以及高、低工作记忆能力者的听力努力。
Ear Hear. 2022;43(4):1164-1177. doi: 10.1097/AUD.0000000000001192. Epub 2022 Jan 4.
10
Weighting of Prosodic and Lexical-Semantic Cues for Emotion Identification in Spectrally Degraded Speech and With Cochlear Implants.频谱减损语音和人工耳蜗语音中韵律和词汇语义线索的加权用于情感识别。
Ear Hear. 2021;42(6):1727-1740. doi: 10.1097/AUD.0000000000001057.

引用本文的文献

1
Fast frequency modulation is encoded according to the listener expectations in the human subcortical auditory pathway.快速频率调制是根据人类皮层下听觉通路中的听众期望进行编码的。
Imaging Neurosci (Camb). 2024 Sep 19;2. doi: 10.1162/imag_a_00292. eCollection 2024.
2
An active inference account of stuttering behavior.口吃行为的主动推理解释。
Front Hum Neurosci. 2025 Apr 3;19:1498423. doi: 10.3389/fnhum.2025.1498423. eCollection 2025.
3
Cochlear implantation in adults with acquired single-sided deafness improves cortical processing and comprehension of speech presented to the non-implanted ears: a longitudinal EEG study.

本文引用的文献

1
Generative models, linguistic communication and active inference.生成模型、语言交流和主动推理。
Neurosci Biobehav Rev. 2020 Nov;118:42-64. doi: 10.1016/j.neubiorev.2020.07.005. Epub 2020 Jul 17.
2
Reduced prediction error responses in high-as compared to low-uncertainty musical contexts.高不确定性音乐环境下的预测误差响应低于低不确定性音乐环境。
Cortex. 2019 Nov;120:181-200. doi: 10.1016/j.cortex.2019.06.010. Epub 2019 Jun 28.
3
Neuronal message passing using Mean-field, Bethe, and Marginal approximations.使用平均场、Bethe 和边缘近似进行神经元信息传递。
成人获得性单侧耳聋患者的人工耳蜗植入可改善对呈现给未植入耳的语音的皮质处理和理解:一项纵向脑电图研究。
Brain Commun. 2025 Jan 3;7(1):fcaf001. doi: 10.1093/braincomms/fcaf001. eCollection 2025.
4
Generative models for sequential dynamics in active inference.主动推理中序列动力学的生成模型。
Cogn Neurodyn. 2024 Dec;18(6):3259-3272. doi: 10.1007/s11571-023-09963-x. Epub 2023 Apr 26.
5
Using the Mismatch Negativity to Evaluate Hearing Aid Directional Enhancement Based on Multistream Architecture.利用失配负波评估基于多流架构的助听器定向增强功能。
Ear Hear. 2025;46(3):747-757. doi: 10.1097/AUD.0000000000001619. Epub 2024 Dec 19.
6
Slow but flexible or fast but rigid? Discrete and continuous processes compared.慢而灵活还是快而刻板?离散与连续过程之比较。
Heliyon. 2024 Oct 18;10(20):e39129. doi: 10.1016/j.heliyon.2024.e39129. eCollection 2024 Oct 30.
7
A Broken Duet: Multistable Dynamics in Dyadic Interactions.一曲破碎的二重奏:二元互动中的多稳态动力学
Entropy (Basel). 2024 Aug 28;26(9):731. doi: 10.3390/e26090731.
8
Sensorimotor learning during synchronous speech is modulated by the acoustics of the other voice.同步言语过程中的感觉运动学习受另一个声音的声学特征调节。
Psychon Bull Rev. 2025 Feb;32(1):306-316. doi: 10.3758/s13423-024-02536-x. Epub 2024 Jul 2.
9
Natural language syntax complies with the free-energy principle.自然语言句法符合自由能原理。
Synthese. 2024;203(5):154. doi: 10.1007/s11229-024-04566-3. Epub 2024 May 3.
10
Federated inference and belief sharing.联邦推理与信念共享。
Neurosci Biobehav Rev. 2024 Jan;156:105500. doi: 10.1016/j.neubiorev.2023.105500. Epub 2023 Dec 5.
Sci Rep. 2019 Feb 13;9(1):1889. doi: 10.1038/s41598-018-38246-3.
4
Advances in Variational Inference.变分推理的进展
IEEE Trans Pattern Anal Mach Intell. 2019 Aug;41(8):2008-2026. doi: 10.1109/TPAMI.2018.2889774. Epub 2018 Dec 25.
5
Cortical Response to the Natural Speech Envelope Correlates with Neuroimaging Evidence of Cognition in Severe Brain Injury.大脑皮层对自然语音包络的反应与严重脑损伤认知的神经影像学证据相关。
Curr Biol. 2018 Dec 3;28(23):3833-3839.e3. doi: 10.1016/j.cub.2018.10.057. Epub 2018 Nov 21.
6
Familiar Voices Are More Intelligible, Even if They Are Not Recognized as Familiar.熟悉的声音更容易理解,即使它们没有被识别为熟悉的声音。
Psychol Sci. 2018 Oct;29(10):1575-1583. doi: 10.1177/0956797618779083. Epub 2018 Aug 10.
7
Predicting language outcomes after stroke: Is structural disconnection a useful predictor?预测脑卒中后的语言预后:结构失连接是否是一个有用的预测指标?
Neuroimage Clin. 2018 Mar 30;19:22-29. doi: 10.1016/j.nicl.2018.03.037. eCollection 2018.
8
Statistical learning and probabilistic prediction in music cognition: mechanisms of stylistic enculturation.音乐认知中的统计学习与概率预测:风格文化适应机制
Ann N Y Acad Sci. 2018 May 11;1423(1):378-95. doi: 10.1111/nyas.13654.
9
Speech Intelligibility Predicted from Neural Entrainment of the Speech Envelope.基于语音包络神经同步预测语音可懂度。
J Assoc Res Otolaryngol. 2018 Apr;19(2):181-191. doi: 10.1007/s10162-018-0654-z. Epub 2018 Feb 20.
10
The graphical brain: Belief propagation and active inference.图形大脑:信念传播与主动推理。
Netw Neurosci. 2017;1(4):381-414. doi: 10.1162/NETN_a_00018. Epub 2017 Dec 31.