Explainable Machine Learning for Benchmarking National Digital Research Ecosystems: Implications for Higher-Education Strategy

Authors

  • Rissa A. Lasap
  • Luigi Carlo M. De Jesus

Keywords:

explainable artificial intelligence; research analytics; higher education; digital infrastructure; national innovation systems; secondary data; random forest

Abstract

National research systems increasingly depend on the joint availability of scientific human capital, sustained research and development (R&D) investment, and trustworthy digital infrastructure. Yet international benchmarking often relies on single indicators or opaque composite rankings that obscure the system conditions associated with research output. This study develops an explainable machine-learning benchmark of national digital research ecosystems and translates the findings into higher-education strategy. Twelve open World Development Indicators were combined in a country-year data architecture. The outcome was 2023 scientific and technical journal articles per million population; predictors were the latest available values from 2019–2023 for R&D intensity, researcher density, tertiary enrollment, economic capacity, high-technology exports, secure servers, public education expenditure, electricity access, mobile subscriptions, and fixed broadband. The merged public dataset contains 197 countries. A primary analytic cohort of 99 countries was defined by observed R&D intensity and researcher density, while limited missingness in the remaining predictors was imputed only within training folds. Four prespecified models were evaluated using five-times repeated five-fold cross-validation. Random forest achieved the lowest held-out error (mean R²=0.820; RMSE=0.613 on the log1p outcome; Spearman ρ=0.912), closely followed by extremely randomized trees. Cross-validated permutation analysis identified researcher density as the dominant predictor, followed by GDP per capita, secure Internet servers, R&D intensity, and electricity access. Complete-case sensitivity analysis (n=79) produced similar results (R²=0.810; ρ=0.914). The results support a capacity-sequencing interpretation: universities cannot treat AI adoption, digital transformation, research workforce development, and basic infrastructure as independent initiatives. The benchmark is descriptive rather than causal and is released with country-level source years, inclusion flags, model diagnostics, a data dictionary, and reproducible analysis files.

Downloads

Published

2026-09-28

How to Cite

Lasap, R. A., & De Jesus, L. C. M. (2026). Explainable Machine Learning for Benchmarking National Digital Research Ecosystems: Implications for Higher-Education Strategy. International Journal of Artificial Intelligence and Machine Learning, 6(12s), 426–435. Retrieved from https://www.svedbergopen.com/index.php/ijaiml/article/view/2428