A fast integrative clustering and feature selection approach for high-dimensional multiview data
Abdalkarim Alnajjar et al.
What the paper says
Cluster analysis has been widely used in biomedical studies for disaggregating heterogeneous diseases and identifying disease subtypes that may inform clinical decisions. In the era of advanced data science and engineering, cluster analysis faces new challenges due to high dimensionality, multimodality and computational complexity. In the present study, we propose a fast integrative clustering approach based on variational Bayesian inference, called iClusterVB. The iClusterVB enables the integration of multiple datasets into the clustering process while performing feature selection in high-dimensional settings for mixed data types, including continuous, categorical, and count data. Simulation studies are performed to compare the performance of iClusterVB with six competing methods and highlight its advantages. Additionally, iClusterVB is applied to three real-life studies to demonstrate its utility in identifying important features and cancer subtypes that are associated with distinct survival probabilities. A user-friendly R package iClusterVB and a tutorial are developed to implement the proposed approach.
Evidence weight
Balanced mode · F 0.40 / M 0.15 / V 0.05 / R 0.40
| F · citation impact | 0.50 × 0.4 = 0.20 |
| M · momentum | 0.50 × 0.15 = 0.07 |
| V · venue signal | 0.50 × 0.05 = 0.03 |
| R · text relevance † | 0.50 × 0.4 = 0.20 |
† Text relevance is estimated at 0.50 on the detail page — for your query’s actual relevance score, open this paper from a search result.