凝胶形成黏蛋白结构域的串联重复结构可通过单分子实时测序数据揭示。

Tandem repeats structure of gel-forming mucin domains could be revealed by SMRT sequencing data.

机构信息

Big Data Decision Institution, Jinan University, No 601 Huangpu Avenue West, Tianhe District, Guangzhou, 510632, China.

出版信息

Sci Rep. 2022 Nov 30;12(1):20652. doi: 10.1038/s41598-022-25262-7.

DOI:10.1038/s41598-022-25262-7

PMID:36450890

原文链接:https://pmc.ncbi.nlm.nih.gov/articles/PMC9712336/

Abstract

Mucins are large glycoproteins that cover and protect epithelial surface of the body. Mucin domains of gel-forming mucins are rich in proline, threonine, and serine that are heavily glycosylated. These domains show great complexity with tandem repeats, thus make it difficult to study the sequences. With the coming of single molecule real-time (SMRT) sequencing technologies, we manage to present sequence structure of mucin domains via SMRT long reads for gel-forming mucins MUC2, MUC5AC, MUC5B and MUC6. Our study shows that for different individuals, single nucleotide polymorphisms could be found in mucin domains of MUC2, MUC5AC, MUC5B and MUC6, while different number of tandem repeats could be found in mucin domains of MUC2 and MUC6. Furthermore, we get the sequence of MUC2, MUC5AC, and MUC5B mucin domain in a Chinese individual for each nucleotide at accuracy of possibly 99.98-99.99%, 99.93-99.99%, and 99.76-99.99%, respectively. We report a new method to obtain DNA sequence of gel-forming mucin domains. This method will provided new insights on getting the sequence for Tandem Repeat parts which locate in coding region. With the sequences we obtained through this method, we can give more information for people to study the sequences of gel-forming mucin domains.

摘要

粘蛋白是覆盖和保护身体上皮表面的大型糖蛋白。形成凝胶的粘蛋白的粘蛋白结构域富含脯氨酸、苏氨酸和丝氨酸，这些氨基酸高度糖基化。这些结构域具有串联重复的特点，因此很难研究其序列。随着单分子实时（SMRT）测序技术的出现，我们成功地通过 SMRT 长读长展示了形成凝胶的粘蛋白 MUC2、MUC5AC、MUC5B 和 MUC6 的粘蛋白结构域的序列结构。我们的研究表明，在不同个体中，MUC2、MUC5AC、MUC5B 和 MUC6 的粘蛋白结构域中可能存在单核苷酸多态性，而 MUC2 和 MUC6 的粘蛋白结构域中可能存在不同数量的串联重复。此外，我们以可能高达 99.98-99.99%、99.93-99.99%和 99.76-99.99%的准确性获得了一个中国人 MUC2、MUC5AC 和 MUC5B 粘蛋白结构域的每个核苷酸的序列。我们报告了一种获得形成凝胶的粘蛋白结构域 DNA 序列的新方法。这种方法将为位于编码区的串联重复部分的序列获取提供新的见解。通过我们使用这种方法获得的序列，我们可以为人们研究形成凝胶的粘蛋白结构域的序列提供更多信息。