P.C. Carvalho, J. Hewel, V.C. Barbosa, J.R. Yates III
Published: April 15, 2008
Genet. Mol. Res. 7 (2) : 342-356
DOI: https://doi.org/10.4238/vol7-2gmr426
Cite this Article:
P.C. Carvalho, J. Hewel, V.C. Barbosa, J.R. Yates (2008). Identifying differences in protein expression levels by spectral counting and feature selection. Genet. Mol. Res. 7(2): 342-356. https://doi.org/10.4238/vol7-2gmr426
About the Authors
P.C. Carvalho, J. Hewel, V.C. Barbosa, J.R. Yates III
Corresponding author
P.C. Carvalho
E-mail: carvalhopc@cos.ufrj.br
ABSTRACT
Spectral counting is a strategy to quantify relative protein concentrations in pre-digested protein mixtures analyzed by liquid chromatography online with tandem mass spectrometry. In the present study, we used combinations of normalization and statistical (feature selection) methods on spectral counting data to verify whether we could pinpoint which and how many proteins were differentially expressed when comparing complex protein mixtures. These combinations were evaluated on real, but controlled, experiments (yeast lysates were spiked with protein markers at different concentrations to simulate differences), which were therefore verifiable. The following normalization methods were applied: total signal, Z-normalization, hybrid normalization, and log preprocessing. The feature selection methods were: the Golub index, the Student t-test, a strategy based on the weighting used in a forward-support vector machine (SVM-F) model, and SVM recursive feature elimination. The results showed that Z-normalization combined with SVM-F correctly identified which and how many protein markers were added to the yeast lysates for all different concentrations. The software we used is available at http://pcarvalho.com/patternlab.
Key words: Feature selection, MudPIT, Support vector machine, Spectral counting, Feature ranking, Support vector machine.