Research article · Applied Artificial Intelligence
Calibrated channel omnibus feature selection for ultra-high-dimensional biomedical data
Why no single relevance criterion is safe when the effect composition is unknown
Authors
Abstract
Feature selection on ultra-high-dimensional biomedical arrays is routinely benchmarked with a single relevance criterion, and comparative studies of methods such as mRMR and SVM-RFE report accuracy without asking whether the criterion matches the way the classes actually differ. This work shows that the question is not incidental. Classes may separate through a shift in location, a shift in dispersion at equal means, or a switch-like activation restricted to a subgroup, and a criterion consistent against one of these families can be no better than chance against another. Because the composition of effects in a real assay is unknown before selection, committing to one criterion is a wager rather than a design choice. We propose CALICO-FS, a single-stage omnibus ranker that maps three complementary dependence statistics onto a common permutation-calibrated scale, estimates the reliability of each from excess tail mass, and fuses them by a reliability-weighted mean. On a controlled benchmark of 12625 features and 21 samples with planted ground truth, CALICO-FS attains a worst-case recall over four effect regimes of 0.288 against 0.198 for the strongest of six baselines, and the best mean recall of 0.443. Across ninety-six paired comparisons it is never significantly worse than any baseline in any regime and is significantly better in twenty of twenty-four. We also report two negative findings that qualify the contribution: the method is less stable under resampling than classical filters, and its recovery advantage does not transfer to downstream classification accuracy.