主页 文献库文献详情
PMID: 41781977 已发表 · epublish 英语

Disease- and gene-specific deep learning for pathogenicity prediction of rare missense variants in cancer predisposition genes.

BioData mining ·第 19 卷 ·第 1 期 ·2026-03-04

Lee DB, Kang HU, Hwang KB

摘要

BACKGROUND: Hereditary cancers frequently arise from germline pathogenic variants, yet only a small proportion of reported variants have been clinically classified, leaving most missense variants unresolved as variants of uncertain significance (VUS). Although recent machine-learning approaches have explored disease-specific or gene-specific contexts to improve pathogenicity prediction, these models remain fundamentally limited by the scarcity of labeled data and the underutilization of abundant VUS. RESULTS: We propose a deep-learning framework that integrates autoencoder pretraining with a deep ensemble strategy to improve variant pathogenicity prediction, effectively leverage unlabeled VUS during pretraining, and reduce uncertainty arising from limited training samples. To validate each component of our framework, we evaluated its performance under both disease-specific and gene-specific training setups. Experiments on ClinVar variants from BRCA1, BRCA2, MLH1, and MSH2 showed that our framework achieved the best performance in the gene-specific setup for BRCA1—likely because BRCA1 contains substantially more gene-specific training data than the other genes—whereas the disease-specific setup yielded superior results for the remaining genes, which had comparatively limited gene-specific samples. Overall, our method significantly outperformed existing approaches. We also introduce an interpretability approach that provides variant-level importance profiles across pathogenicity classes, thereby enhancing transparency and clinical applicability. Moreover, by projecting feature-level importance scores into a two-dimensional space, we demonstrate that pretraining enables the model to learn distinctly different feature representations, illustrating how pretraining and ensemble learning synergistically contribute to improved predictive performance. CONCLUSIONS: Our framework preserves the specificity of disease- and gene-specific approaches, overcomes data scarcity through VUS-guided pretraining and ensembling, and offers interpretable outcomes that may be helpful for clinical decision support. Moreover, our results suggest a promising direction for pathogenicity prediction of rare missense variants and indicate that the proposed framework may be extendable to additional genes under appropriate data and modeling conditions.

关键词
Deep ensemble Deep neural networks Denoising autoencoders Predisposition cancer genes Rare missense variant Representation learning Variant pathogenicity prediction
文献信息
期刊
BioData mining
期刊简称
BioData Min
ISSN
1756-0381
发表日期
2026-03-04
语言
英语
国家/地区
England
NLM ID
101319161
分析服务
分析服务

联系地址

山东省济南市章丘区文博路2号

齐鲁师范学院 genelibs生信实验室

山东省济南市高新区舜华路750号

大学科技园北区F座4单元2楼

电话: 0531-88819269

微信公众号

关注微信订阅号,实时查看信息,关注医学生物学动态。


商务邮箱

E-mail: product@genelibs.com