高级检索

    基于文本-图模态的癌症驱动基因识别方法

    Text-Image Modal-Based Method for Identifying Cancer Driver Genes

    • 摘要: 癌症驱动基因的识别是肿瘤学研究中的一个重要方向,对于深入解析癌症的发生机制及实现精准医疗具有重要价值。目前,大多数研究方法主要依赖图数据来捕捉驱动基因的特征,但是癌症的发生过程涉及多种生物分子之间复杂的相互作用,仅凭单一模态信息难以全面揭示其本质。因此,如何高效整合多模态生物数据以全面刻画驱动基因,成为亟待解决的关键问题。为此,本研究提出了一种基于文本-图模态的癌症驱动基因识别方法TGMNN(text-graph multi-modal neural network)。该方法在图模态的基础上,进一步引入与基因相关的文本语义信息,并采用交互注意力机制实现多模态的协同融合,从而降低模态之间的差异。与现有方法相比,TGMNN能够更有效地捕捉多模态之间的信息交互,显著提升驱动基因预测的准确性。实验结果表明,该方法的AUC值达到0.897,优于现有方法,突显了模型的有效性。

       

      Abstract: The identification of cancer driver genes is an important research area in oncology research, holding significant value for a deeper understanding of the mechanisms of cancer development and achieving precision medicine. Currently, most research methods primarily rely on graph data to capture the characteristics of driver genes. However, the evolution process of cancer development involves many complex interactions among various biomolecules, making it difficult to fully reveal its essence using information from only single modality. Therefore, how to efficiently integrate multi-modal biological data to comprehensively identify driver genes has become a critical issue that needs to be addressed. To this end, this study proposes a cancer driver gene identification method, titled as TGMNN (text-graph multi-modal neural network), based on a combination of text and graph modality. This method builds upon the graph modality by further incorporating text semantic information related to genes and employs an interactive attention mechanism to achieve collaborative fusion of multi-modal data, thereby reducing the differences between modalities. Compared to existing methods, TGMNN can more effectively capture the information interactions between modalities, significantly improving the accuracy of driver gene predictions. Experimental results show that the AUC value of this method reaches 0.897, outperforming existing methods and high-lighting the model's effectiveness.

       

    /

    返回文章
    返回