登录 EN

添加临时用户

基于生物序列与结构大数据的分析方法及应用研究

New Methods and Applications Research on Big Data of Biological Sequence and Structure

作者:赵鑫
  • 学号
    2015******
  • 学位
    博士
  • 电子邮箱
    zha******.cn
  • 答辩日期
    2020.05.19
  • 导师
    丘成栋
  • 学科名
    数学
  • 页码
    110
  • 保密级别
    公开
  • 培养单位
    042 数学系
  • 中文关键词
    序列分析,分类及系统发育学分析,改进的自然向量法,凸分析原理,蛋白质结构比较
  • 英文关键词
    Sequence Analysis, Classification and Phylogenetic Analysis, Improved Natural Vector Method, Convex Hull Principle, Protein Structure Comparison

摘要

随着科学技术的发展和进步,生物序列及结构数据在以指数增长的速度增加,如何高效地从海量数据中提取关键信息,预测新数据所属类别,以及根据序列与结构数据表现的相似性从而研究物种间的系统发育关系成为近年来生物信息学中重要的研究问题。生物信息学中传统的序列比对以及结构比对方法通常复杂度较高、运算时间久,难以处理数据规模大或是结构较复杂数据,这就启发我们提出序列分析以及结构分析的新方法。在对序列分析的研究中,为了进一步研究序列中核苷酸或氨基酸的分布情况,本文提出核苷酸(氨基酸)之间的相关性概念,并将此特征其加入到原始的自然向量中,提出改进的自然向量法。它是一种非比对的序列表示法,将每条序列转化为一个向量,并且每条序列和它所对应的自然向量存在一一对应关系。此算法计算复杂度低,能够准确有效地反映序列信息,通过向量之间的距离实现序列相似性的比较,是研究物种分类学以及系统发育学关系的有力工具。此外,本文在研究中将自助抽样法与改进的自然向量方法结合,计算构建的系统发生树的置信概率。新方法在多组细菌、真菌、病毒等大规模数据集的具体应用研究中表现出了很强的鲁棒性。研究将新方法与多种序列分析的方法比较,以及根据系统发育关系的置信性检验,验证新方法的可靠性与稳定性。为了进一步探究相同物种或同一家族包含序列的自然向量在空间中的分布规律,本文提出凸分析原理,并且在包括人类全部蛋白质数据以及真核生物蛋白激酶在内的多个数据规模大、物种(家族)类别个数多的可靠数据集上验证了凸分析原理,即不同物种(家族)包含序列对应的自然向量构成凸包是互不相交的。基于凸分析原理,每个物种可以由自然向量构建的凸包中心点来表示,不同物种的相似性通过中心点之间的欧氏距离来衡量,凸分析方法结合自然向量方法为研究分子生物学中的物种间系统发育关系提供了新视角,也为在给定类别中寻找新基因或蛋白质数据提供了新手段。

In recent years, classifying and analyzing biological big data has become one of the most important research areas in bioinformatics as the tools for getting biological sequences and structures increase. Obtaining information for different kinds of data by reasonable and effective methods, comparing the similarity or dissimilarity of sequences or structures, analyzing the evolutionary relationship of different species, and determining the class of problematic data have become very important research directions. In addition, predicting the class of new data and inferring their biological functions or properties are also important research problems in taxonomy. Traditional taxonomic methods are usually too complex to make it difficult to deal with the classification and phylogenetic analysis of a large number of sequences or data with complicated structures. This motivates us to propose new methods for sequence and structure analysis.In this dissertation, we will introduce the improved natural vector method for sequence analysis. In order to further study the distribution of nucleotides and amino acids in one sequence, we first purpose the correlation of nucleotides and amino acids and add this feature to the traditional natural vector. It is a non-aligned rapid representation for sequences.Each sequence is converted into a vector by extracting information such as the number, average position and central moment of each nucleotide or amino acid in the sequence. The correspondence between the sequence and its natural vector is one-to-one. This algorithm can reflect sequence information accurately and effectively, and complete sequence comparison based on the distances between vectors with low computational complexity. It is a powerful tool for classification and phylogenetic analysis.We combine the bootstrapping method and natural vector method to calculate the confidence probabilities on phylogenetic trees. Using new method, we systematically analyze four data sets of alphaproteobacterial proteomes in order to reconstruct the phylogeny of Alphaproteobacteria. In addition, this method is also used to resolve some evolutionary relationships of Prochlorococcus, that the strains SS120 and MIT9211 do not form a monophyletic clade in the phylogeny of Prochlorococcus. Furthermore, The classification performs well on the fungi barcode dataset with high and robust accuracy. The reasonable phylogenetic trees of $\beta$-coronaviruses we obtained further validate the new method.In order to further explore the distribution law of natural vectors for the sequences in high dimensional space, we verify that the convex hulls constructed by the natural vectors of sequences from the same family do not intersect with convex hulls of natural vectors from other families, and propose the convex hull principle. This principle indicates that sequences with similar distribution should be in the same family. We verify this principle computationally by using all available and reliable sequences on protein kinase datasets and human proteins as well as all DNA barcodes. It also provides a quick way to search for natural vector points that lie within the convex hull of a given class and discover new sequence. This will open up a new interdisciplinary research in biology, mathematics and computer science.For structure comparison, we extend the Yau-Hausdorff method, which achieves the best match of two protein structures accurately in a fast way with low complexity. This method measures the similarity or dissimilarity of structures on account of descending dimension in calculation without losing any information. It can also infer protein function by structural similarity. The new algorithm is compared with some traditional structure comparison methods to show the accuracy and stability of our approach in structure comparison.