Home LiteratureArticle Details
PMID: 33410688 Published · ppublish English Journal Article Research Support, Non-U.S. Gov't

TargetDBP+: Enhancing the Performance of Identifying DNA-Binding Proteins via Weighted Convolutional Features.

Journal of chemical information and modeling ·Vol. 61 ·No. 1 ·2021-00-25 ·页码 505-515

Hu J, Rao L, Zhu YH, Zhang GJ, Yu DJ

Abstract

Protein-DNA interactions exist ubiquitously and play important roles in the life cycles of living cells. The accurate identification of DNA-binding proteins (DBPs) is one of the key steps to understand the mechanisms of protein-DNA interactions. Although many DBP identification methods have been proposed, the current performance is still unsatisfactory. In this study, a new method, called TargetDBP+, is developed to further enhance the performance of identifying DBPs. In TargetDBP+, five convolutional features are first extracted from five feature sources, i.e., amino acid one-hot matrix (AAOHM), position-specific scoring matrix (PSSM), predicted secondary structure probability matrix (PSSPM), predicted solvent accessibility probability matrix (PSAPM), and predicted probabilities of DNA-binding sites (PPDBSs); second, the five features are weightedly and serially combined using the weights of all of the elements learned by the differential evolution algorithm; and finally, the DBP identification model of TargetDBP+ is trained using the support vector machine (SVM) algorithm. To evaluate the developed TargetDBP+ and compare it with other existing methods, a new gold-standard benchmark data set, called UniSwiss, is constructed, which consists of 4881 DBPs and 4881 non-DBPs extracted from the UniprotKB/Swiss-Prot database. Experimental results demonstrate that TargetDBP+ can obtain an accuracy of 85.83% and precision of 88.45% covering 82.41% of all DBP data on the independent validation subset of UniSwiss, with the MCC value (0.718) being significantly higher than those of other state-of-the-art control methods. The web server of TargetDBP+ is accessible at http://csbio.njust.edu.cn/bioinf/targetdbpplus/; the UniSwiss data set and stand-alone program of TargetDBP+ are accessible at https://github.com/jun-csbio/TargetDBPplus.

MeSH 主题词
Algorithms Binding Sites DNA-Binding Proteins/metabolism Databases, Protein Position-Specific Scoring Matrices Support Vector Machine
化学物质
DNA-Binding Proteins
作者与单位
共 5 位作者,点击展开单位 / ORCID
Hu Jun ORCID
College of Information Engineering, Zhejiang University of Technology, Hangzhou 310023, P. R. China. | Key Laboratory of Data Science and Intelligence Application, Fujian Province University, Zhangzhou 363000, P. R. China.
Rao Liang
College of Information Engineering, Zhejiang University of Technology, Hangzhou 310023, P. R. China.
Zhu Yi-Heng
School of Computer Science and Engineering, Nanjing University of Science and Technology, Xiaolingwei 200, Nanjing 210094, P. R. China.
Zhang Gui-Jun
College of Information Engineering, Zhejiang University of Technology, Hangzhou 310023, P. R. China.
Yu Dong-Jun ORCID
School of Computer Science and Engineering, Nanjing University of Science and Technology, Xiaolingwei 200, Nanjing 210094, P. R. China.
Article Info
Journal
Journal of chemical information and modeling
Abbr.
J Chem Inf Model
ISSN
1549-960X
Published
2021-00-25
电子出版
2021-00-07
页码
505-515
Language
English
Country/Region
United States
NLM ID
101230060
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com