A Mathematical and Statistical Framework For Uncertainty-Aware Machine Learning in Complex Data Environments
Keywords:
uncertainty-aware machine learning, high-dimensional data, predictive entropy, probability calibration, selective predictionAbstract
High-dimensional data sets can be used to train machine-learning models that can show high discrimination abilities, but also can give probability estimates with varying calibrations and reliabilities. This study developed and evaluated a mathematical and statistical framework integrating predictive performance, calibration, uncertainty, selective prediction, and predictor stability in a complex high-dimensional classification setting. A total of 174 observations and 450 numerical predictors were included in the analysis. To assess model performance, repeated stratified five-fold cross-validation with 10 repetitions was applied to Logistic Regression, Support Vector Machine, Random Forest, and Gradient Boosting. The Random Forest model had the best ROC-AUC (0.961) and accuracy (0.880), while the Gradient Boosting model had the lowest Brier score (0.0925) and negative log-loss (0.3188). Support Vector Machine had the lowest value of ECE (0.0450). Predictive entropy was higher for misclassifications across all predictors and had the strongest correlation with participant-level misclassification for Gradient Boosting (rho = 0.7065). When using uncertainty-based selective prediction, the accuracy reached 1.000 for 50% coverage for Random Forest and Gradient Boosting, respectively. The predictor effects also showed good consistency between cross-validation folds. The results suggest a more complete picture of the reliability of machine learning models in high-dimensional data settings can be obtained by considering discrimination, probabilistic reliability, predictive uncertainty, selective coverage, and predictor stability together.





