Thurnhofer-Hemsi Karl, García-Aguilar Iván, Fernández-Rodriguez José David, López-Rubio Ezequiel
ITIS Software, Universidad de Málaga, C/Arquitecto Francisco Peñalosa 18, Málaga 29010, Spain.
Biomedic Research Institute of Málaga, IBIMA Plataforma BIONAND, C/Doctor Miguel Díaz Recio, 28, Málaga 29010, Spain.
J Chem Inf Model. 2025 Aug 11;65(15):7936-7955. doi: 10.1021/acs.jcim.5c00354. Epub 2025 Jul 28.
The processing of chemical information by computational intelligence methods faces the challenge of the structural complexity of molecular graphs. These graphs are not amenable to being represented in a suitable way for such methods. The most popular representation is the SMILES notation standard. However, it comes with some limitations, such as the abundance of nonvalid strings and the fact that similar strings often represent very different molecules. In this work, a completely different approach to chemical nomenclature is presented. A reduced instruction set is defined, and the language of all strings that are sequences of such instructions is considered. Instructions provide the means to incrementally add atoms and modify the connectivity of the chemical bonds of atoms to be inserted. Instructions are carefully crafted to guarantee that all strings of this language are valid, i.e., each string represents a molecule. Moreover, slight changes in a string usually correspond to small modifications in the represented molecule. Therefore, this approach is appropriate for use in state-of-the-art computational intelligence systems for chemical information processing, including deep learning models.
利用计算智能方法处理化学信息面临着分子图结构复杂性的挑战。这些图难以以适合此类方法的方式进行表示。最流行的表示方式是SMILES符号标准。然而,它存在一些局限性,例如存在大量无效字符串,并且相似的字符串往往代表非常不同的分子。在这项工作中,提出了一种完全不同的化学命名方法。定义了一个精简指令集,并考虑了所有由这些指令序列组成的字符串的语言。指令提供了逐步添加原子并修改要插入原子的化学键连接性的方法。指令经过精心设计,以确保该语言的所有字符串都是有效的,即每个字符串都代表一个分子。此外,字符串的微小变化通常对应于所表示分子的微小修改。因此,这种方法适用于用于化学信息处理的先进计算智能系统,包括深度学习模型。