A Machine Learning-Based Framework for Predictive and Continuous Cybersecurity GRC Automation: Aligning with Saudi Arabia's NCA Standards
Keywords:
cybersecurity governance; risk and compliance automation; machine learning; National Cybersecurity Authority (NCA); regulatory technology (RegTech); anomaly detection; compliance forecasting.Abstract
Periodic, manual audit processes that are out of step with the pace of contemporary threats and configuration changes continue to dominate cybersecurity governance, risk, and compliance (GRC) practice. This gap is particularly problematic for Saudi Arabian organizations governed by the National Cybersecurity Authority (NCA) framework, which includes the Cybersecurity Maturity Model (CMM), Cloud Cybersecurity Controls (CCC), Operational Technology Cybersecurity Controls (OTCC), and Essential Cybersecurity Controls (ECC). A five-layer machine learning (ML) framework that (i) parses NCA regulatory text into machine-readable control requirements, (ii) normalizes multi-source organizational telemetry into a unified schema, (iii) applies a multi-model ML engine combining gradient-boosted classification, unsupervised drift detection, and multi-horizon time-series forecasting, (iv) presents results through a compliance intelligence dashboard, and (v) continuously improves itself through a federated feedback loop. The compliance-status classifier's macro F1-score is 0.828, and the macro-AUC is 0.948 (with a five-fold cross-validated F1 score of 0.815 ± 0.009) according to the original XGBoost model; the subsequent baseline comparison showed that the classifier based on Logistic Regression is superior. LogReg achieved an increment in macro F1 to 0.858 and macro AUC to 0.963. Based on these results, LogReg was selected as the main Compliance Status Classifier. The drift detector was reported to have reached an AUC of 0.811 and had its corresponding rate decreased from 63.2% to 26.2%. The risk forecaster's overall mean absolute error is in the range of 0.044–0.101 for 7-, 14-, and 30-day outcomes and the continuous pipeline provided a decrease in expenditure for detecting compliance drift by 97.5% compared to semi-annual audits. These results show a residual precision–recall trade-off for boundary cases between partially compliant and non-compliant states, which encourages future work on richer, real-world-calibrated data. Because the ground-truth compliance labels are themselves derived from the same twelve telemetry features used as model inputs, and no independently collected or expert-labelled organizational data were used, these results are reported as a proof-of-concept demonstration of the analytical pipeline rather than as validated real-world performance. They also quantitatively demonstrate that predictive, telemetry-driven GRC automation can significantly outperform static audit cycles under this proof-of-concept, simulated evaluation. The framework and its evaluation protocol offer a reproducible template for regulatory technology (RegTech) research in Saudi Arabia and other jurisdictions with codified, control-based cybersecurity mandates. The framework's analytical core—compliance categorization, drift detection, and risk forecasting—is covered by the empirical evaluation shown here; the regulatory-parsing and telemetry-normalization layers are described architecturally but were not independently benchmarked in this study.





