《儿童与青少年书籍词汇表》(CYP-LEX):一个大规模的词汇数据库,收录了英国儿童和青少年阅读的书籍。
The Children and Young People's Books Lexicon (CYP-LEX): A large-scale lexical database of books read by children and young people in the United Kingdom.
机构信息
Department of Psychology, Royal Holloway, University of London, Egham, UK.
Department of Psychology, University of Milano-Bicocca, Milan, Italy.
出版信息
Q J Exp Psychol (Hove). 2024 Dec;77(12):2418-2438. doi: 10.1177/17470218241229694. Epub 2024 Mar 12.
This article introduces the Children and Young People's Books-Lexicon (CYP-LEX), a large-scale lexical database derived from books popular with children and young people in the United Kingdom. CYP-LEX includes 1,200 books evenly distributed across three age bands (7-9, 10-12, 13+) and comprises over 70 million tokens and over 105,000 types. For each word in each age band, we provide its raw and Zipf-transformed frequencies, all parts-of-speech in which it occurs with raw frequency and lemma for each occurrence, and measures of count-based contextual diversity. Together and individually, the three CYP-LEX age bands contain substantially more words than any other publicly available database of books for primary and secondary school children. Most of these words are very low in frequency, and a substantial proportion of the words in each age band do not occur on British television. Although the three age bands share some very frequent words, they differ substantially regarding words that occur less frequently, and this pattern also holds at the level of individual books. Initial analyses of CYP-LEX illustrate why independent reading constitutes a challenge for children and young people, and they also underscore the importance of reading widely for the development of reading expertise. Overall, CYP-LEX provides unprecedented information into the nature of vocabulary in books that British children aged 7+ read, and is a highly valuable resource for those studying reading and language development.
本文介绍了儿童与青少年书籍词汇库(CYP-LEX),这是一个源自英国儿童和青少年喜爱的书籍的大规模词汇数据库。CYP-LEX 包含 1200 本书,均匀分布在三个年龄段(7-9 岁、10-12 岁和 13+岁),包含超过 7000 万个标记和超过 105000 个词类。对于每个年龄段的每个单词,我们提供其原始和 Zipf 转换频率、以原始频率出现的所有词性以及基于计数的上下文多样性的度量。三个 CYP-LEX 年龄段的词汇量加起来比其他任何为小学和中学儿童提供的公开书籍数据库都要多得多。这些单词中的大多数频率都非常低,而且每个年龄段的单词中都有很大一部分不会出现在英国电视上。尽管这三个年龄段有一些非常常见的单词,但它们在不太常见的单词方面有很大的不同,这种模式在个别书籍中也适用。对 CYP-LEX 的初步分析说明了为什么独立阅读对儿童和青少年来说是一个挑战,也强调了广泛阅读对于阅读专业知识发展的重要性。总体而言,CYP-LEX 前所未有地提供了英国 7 岁以上儿童阅读书籍中词汇本质的信息,是研究阅读和语言发展的人的宝贵资源。