作为数字文化遗产的巴厘语语音数据集。

Balinese text-to-speech dataset as digital cultural heritage.

作者信息

Kadyanan I Gusti Agung Gede Arya, Er Ngurah Agus Sanjaya, Karyawati Anak Agung Istri Ngurah Eka, Putra I Gede Ngurah Arya Wira, Gunawan I Made Suma, Budiantari Ni Made Julia, Octavia Hana Christine

机构信息

Department of Informatics, Faculty of Mathematics and Natural Sciences, Udayana University, Badung, Indonesia.

出版信息

Data Brief. 2025 Apr 9;60:111528. doi: 10.1016/j.dib.2025.111528. eCollection 2025 Jun.

DOI:10.1016/j.dib.2025.111528

PMID:40275973

原文链接:https://pmc.ncbi.nlm.nih.gov/articles/PMC12020862/

Abstract

Balinese language has a complex and unique language level system, yet still lacks representation in speech-based technologies such as Text-to-Speech (TTS) and speech recognition. As one of the linguistically rich regional languages, Balinese language digitization efforts have not been optimally developed, limiting research in natural language processing (NLP) as well as the application of regional language-based voice technologies. The limitation of voice-based datasets in Balinese is a major challenge in the development of this technology. Therefore, this research aims to develop a dataset of Balinese native speaker audio recordings covering various language levels to support applications in Text-to-Speech (TTS) systems, speech recognition, and voice-to-text technology. The dataset was developed through a data acquisition process that involved recording the voices of native Balinese speakers of the Badung dialect. Data was collected by recording the voices of native Balinese speakers using the Badung dialect. The resulting recordings were then processed using denoising techniques to improve audio quality, before being categorized based on Balinese politeness levels (Alus Singgih, Alus Sor, Alus Mider, Mider, and Andap) as well as including additional phrases and alphabets to provide a wider variety to the dataset. The results show that this dataset consists of 1187 recordings that reflect a wide range of social variation in Balinese. By providing this resource, this research not only contributes to the development of speech-based technologies, but also plays a role in the preservation of Balinese in the digital age, as well as opening up further research opportunities in NLP for languages with limited resources.

摘要

巴厘语拥有复杂而独特的语言层级系统，但在诸如文本转语音（TTS）和语音识别等基于语音的技术中仍缺乏代表性。作为语言丰富的地区语言之一，巴厘语的数字化工作尚未得到充分发展，限制了自然语言处理（NLP）研究以及基于地区语言的语音技术应用。巴厘语中基于语音的数据集的局限性是该技术发展的一大挑战。因此，本研究旨在开发一个涵盖各种语言层级的巴厘语母语者音频记录数据集，以支持文本转语音（TTS）系统、语音识别和语音转文本技术中的应用。该数据集是通过数据采集过程开发的，该过程涉及录制巴东方言的巴厘语母语者的声音。通过使用巴东方言录制巴厘语母语者的声音来收集数据。然后，对得到的录音使用去噪技术进行处理以提高音频质量，再根据巴厘语的礼貌程度（文雅庄重、文雅适度、文雅温和、普通、随意）进行分类，并纳入额外的短语和字母，以使数据集更加多样化。结果表明，该数据集由1187条录音组成，反映了巴厘语广泛的社会差异。通过提供这一资源，本研究不仅有助于基于语音的技术发展，还在数字时代巴厘语的保护中发挥作用，同时为资源有限的语言在NLP领域开辟了进一步的研究机会。

https://cdn.ncbi.nlm.nih.gov/pmc/blobs/c0ac/12020862/d0801eeca2f2/gr1.jpg

相似文献

Balinese text-to-speech dataset as digital cultural heritage.作为数字文化遗产的巴厘语语音数据集。

Data Brief. 2025 Apr 9;60:111528. doi: 10.1016/j.dib.2025.111528. eCollection 2025 Jun.

Balinese story texts dataset for narrative text analyses.用于叙事文本分析的巴厘岛故事文本数据集。

Data Brief. 2024 Aug 8;56:110781. doi: 10.1016/j.dib.2024.110781. eCollection 2024 Oct.

A Dataset of Real and Synthetic Speech in Ukrainian.一个乌克兰语真实与合成语音数据集。

Sci Data. 2025 May 6;12(1):745. doi: 10.1038/s41597-025-05084-8.

BanglaSER: A speech emotion recognition dataset for the Bangla language.孟加拉语SER：一个用于孟加拉语的语音情感识别数据集。

Data Brief. 2022 Mar 22;42:108091. doi: 10.1016/j.dib.2022.108091. eCollection 2022 Jun.

LUMINA: Linguistic unified multimodal Indonesian natural audio-visual dataset.LUMINA：印尼语语言统一多模态自然视听数据集。

Data Brief. 2024 Mar 1;54:110279. doi: 10.1016/j.dib.2024.110279. eCollection 2024 Jun.

Video dataset of Balinese dance basic movement for action recognition.用于动作识别的巴厘岛舞蹈基本动作视频数据集。

Data Brief. 2024 Feb 13;53:110189. doi: 10.1016/j.dib.2024.110189. eCollection 2024 Apr.

Do you like my voice? Stakeholder perspectives about the acceptability of synthetic child voices in three South African languages.你喜欢我的声音吗？利益相关者对三种南非语言中合成儿童声音可接受性的看法。

Int J Lang Commun Disord. 2025 Jan-Feb;60(1):e13152. doi: 10.1111/1460-6984.13152.

Do some languages sound more beautiful than others?有些语言听起来比其他语言更优美吗？

Proc Natl Acad Sci U S A. 2023 Apr 25;120(17):e2218367120. doi: 10.1073/pnas.2218367120. Epub 2023 Apr 17.

YembaTones: A syllable-tone annotated dataset for speech recognition and prosodic analysis of the Yemba language.延巴音调：一个用于延巴语语音识别和韵律分析的音节音调标注数据集。

Data Brief. 2023 Nov 27;52:109860. doi: 10.1016/j.dib.2023.109860. eCollection 2024 Feb.

Development of Hausa dataset a baseline for speech recognition.豪萨语数据集的开发——语音识别的一个基线。

Data Brief. 2022 Jan 10;40:107820. doi: 10.1016/j.dib.2022.107820. eCollection 2022 Feb.

本文引用的文献

Balinese story texts dataset for narrative text analyses.用于叙事文本分析的巴厘岛故事文本数据集。

Data Brief. 2024 Aug 8;56:110781. doi: 10.1016/j.dib.2024.110781. eCollection 2024 Oct.

DeepLontar dataset for handwritten Balinese character detection and syllable recognition on Lontar manuscript.用于在 lontar 手稿上手写巴厘文字符检测和音节识别的 DeepLontar 数据集。

Sci Data. 2022 Dec 10;9(1):761. doi: 10.1038/s41597-022-01867-5.

Development of Hausa dataset a baseline for speech recognition.豪萨语数据集的开发——语音识别的一个基线。

Data Brief. 2022 Jan 10;40:107820. doi: 10.1016/j.dib.2022.107820. eCollection 2022 Feb.

文献检索

告别复杂PubMed语法，用中文像聊天一样搜索，搜遍4000万医学文献。AI智能推荐，让科研检索更轻松。

立即免费搜索

文件翻译

保留排版，准确专业，支持PDF/Word/PPT等文件格式，支持 12+语言互译。

免费翻译文档

深度研究

AI帮你快速写综述，25分钟生成高质量综述，智能提取关键信息，辅助科研写作。

立即免费体验

作为数字文化遗产的巴厘语语音数据集。

Balinese text-to-speech dataset as digital cultural heritage.

作者信息

机构信息

出版信息

相似文献

本文引用的文献

文献检索

文件翻译

深度研究

Suppr 超能文献

相似文献

本文引用的文献