A ROBUST MULTI-METHOD FEATURE SELECTION FRAMEWORK FOR EARLY SOFTWARE DEFECT PREDICTION: CROSS-DATASET EVIDENCE FROM PROMISE REPOSITORY
DOI:
https://doi.org/10.4238/vdvsx082Keywords:
Software Defect Prediction; Feature Selection; Feature Stability; Class Imbalance; Software Metrics; Cross Dataset PredictionAbstract
The software defect prediction problem is difficult because of high-dimensional feature spaces, redundancy in software metrics and severe class imbalance. There has been a large body of work proposed on machine learning methods, however, few works have focused on the stability and generalizability of selected features across heterogeneous datasets. A comprehensive multi-method feature selection framework that combines filter (Mutual Information, Chi-Square), wrapper (Recursive Feature Elimination) and embedded (Random Forest) methods and finally a consensus-based ranking mechanism for finding stable and dataset-independent software metrics is proposed. The framework is tested on several PROMISE datasets (KC1, KC2, PC1, JM1), representing various software systems and distributions of metrics. Experimental results show that the size and complexity related metrics such as Lines of Code (LOC), Halstead measures and cyclomatic complexity invariably sit on top of all data sets and selection approaches, resulting in high feature stability. The results of the classification based on Logis-tic Regression, however, have found a significant difference between the measure of precision (0.58-0.65) and the measure of re-call (0.23-0.31), which means that the classification is not so effective in detecting the minority of defective in-stances. The results show that, in imbalanced conditions, using robust feature selection is not enough for early defect prediction. The study highlights the importance of embedding imbalance-aware learning techniques (cost-sensitive learning, synthetic sampling, and ensemble models). The proposed framework shows good feature stability and has a high F1-score in multiple datasets with a value up to 0.42, which means it is effective at cross-dataset defect prediction
Downloads
Published
Issue
Section
License

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.

