主页 文献库文献详情
PMID: 11911798 已发表 · ppublish 英语

Strong feature sets from small samples.

Kim Seungchan, Dougherty Edward R, Barrera Junior, Chen Yidong, Bittner Michael L, Trent Jeffrey M

摘要

For small samples, classifier design algorithms typically suffer from overfitting. Given a set of features, a classifier must be designed and its error estimated. For small samples, an error estimator may be unbiased but, owing to a large variance, often give very optimistic estimates. This paper proposes mitigating the small-sample problem by designing classifiers from a probability distribution resulting from spreading the mass of the sample points to make classification more difficult, while maintaining sample geometry. The algorithm is parameterized by the variance of the spreading distribution. By increasing the spread, the algorithm finds gene sets whose classification accuracy remains strong relative to greater spreading of the sample. The error gives a measure of the strength of the feature set as a function of the spread. The algorithm yields feature sets that can distinguish the two classes, not only for the sample data, but for distributions spread beyond the sample data. For linear classifiers, the topic of the present paper, the classifiers are derived analytically from the model, thereby providing an enormous savings in computation time. The algorithm is applied to cancer classification via cDNA microarrays. In particular, the genes BRCA1 and BRCA2 are associated with a hereditary disposition to breast cancer, and the algorithm is used to find gene sets whose expressions can be used to classify BRCA1 and BRCA2 tumors.

文献信息
期刊
Journal of computational biology : a journal of computational molecular cell biology
期刊简称
J Comput Biol
发表日期
2002-06-24
收录日期
2002-03-25
更新日期
2006-11-15
语言
英语
国家/地区
United States
NLM ID
9433358
分析服务
分析服务

联系地址

山东省济南市章丘区文博路2号

齐鲁师范学院 genelibs生信实验室

山东省济南市高新区舜华路750号

大学科技园北区F座4单元2楼

电话: 0531-88819269

微信公众号

关注微信订阅号,实时查看信息,关注医学生物学动态。


商务邮箱

E-mail: product@genelibs.com