A MULTI-STAGE FEATURE SELECTION AND HYBRID CNN–TRANSFORMER FRAMEWORK WITH BLOCKCHAIN-ENABLED CONSENT GOVERNANCE FOR EXPLAINABLE CLINICAL GENOMIC RISK ASSESSMENT

Authors

  • Kashyap Dave Author
  • Dr. Hitesh kumar M. Nimbark Author

DOI:

https://doi.org/10.4238/43qc4049

Keywords:

Clinical genomic risk assessment; feature selection; mRMR; LASSO; hybrid CNN–Transformer; explainable artificial intelligence; SHAP; blockchain consent management; generalisation gap; high-dimensional low-sample-size data.

Abstract

High-throughput sequencing has made clinical genomic data abundant, but it has not made genomic artificial intelligence trustworthy. Gene-expression matrices are extremely wide and extremely shallow — tens of thousands of transcripts measured over a few hundred patients — a regime in which classifiers memorise rather than generalise, operate as opaque decision engines, and are trained on some of the most re-identifiable data a person can produce. This paper presents a complete methodological progression from an empirically grounded diagnosis of that failure to a governance-aware architecture designed to correct it. A published baseline pipeline was independently reproduced on the GSE68086 tumour-educated-platelet RNA-sequencing benchmark (57,736 transcripts; 173 samples; six cancer classes) using Support Vector Classification, Decision Tree, Multi-Layer Perceptron and one-dimensional Convolutional Neural Network models. All four classifiers exhibited severe train–test divergence: generalisation gaps of 18.3, 45.3, 41.3 and 37.6 percentage points respectively, with the best test accuracy reaching only 60.4 per cent despite 98.1 per cent training accuracy and a multi-class AUC of 0.89. Diagnostic analysis attributed this divergence to three specific and correctable design faults rather than to dataset difficulty alone: unsupervised Shannon-entropy gene ranking that optimises variability instead of class separability, synthetic minority oversampling applied before rather than within the train–test partition, and a purely convolutional inductive bias that cannot model the long-range gene–gene dependencies characteristic of co-regulated transcriptional programmes. Each fault is then addressed explicitly. A three-stage supervised gene selection pipeline is formalised — univariate ANOVA F-test or mutual-information screening, minimum Redundancy-Maximum-Relevance filtering, and cross-validated LASSO confirmation — and is selected through a structured comparison of filter, wrapper, embedded and hybrid families against redundancy handling, small sample stability, computational cost and reproducibility. A hybrid one-dimensional CNN–Transformer classifier is then justified against four competing architectures, coupling convolutional local motif extraction with multi head self-attention over the selected gene tokens. The predictive core is embedded in a layered deployable framework that adds SHAP-based and attention-based explanation, a blockchain smart-contract layer providing dynamic patient consent and immutable access auditing, and a clinician-facing decision-support interface. A systematic review of fifteen studies published in indexed journals during 2025–2026 confirms that no existing work simultaneously compares multiple feature-selection strategies, models both local and global gene interactions, exposes a structured explanation layer, and enforces consent governance. The contribution is therefore an evidence-linked architectural specification with a pre-registered progressive validation protocol, in which every component is traceable to a measured baseline failure.

Downloads

Published

2026-08-05

Issue

Section

Articles