Learned Zero-Day Intrusion Detection: A Systematic Review of Methods, Evaluation Practice and Open Problems, 2020–2026
Keywords:
Anomaly detection, autoencoders, deep learning, evaluation methodology, intrusion detection systems, large language models, network security, reinforcement learning, semi-supervised learning, systematic review, zero-day attack detection.Abstract
Zero-day intrusions are, by construction, attacks for which no signature, patch or prior labelled example exists, so the defensive burden falls on models that generalise from known behaviour to unknown behaviour. The resulting literature has grown quickly and unevenly, and it is now difficult to see which architectural families are actually distinct, which evaluation practices are shared, and which reported gains are attributable to method rather than to protocol. This article systematically reviews forty-seven studies published between 2020 and 2026, drawn from a curated domain corpus of fifty-one records after four duplicate entries were removed. Each study is coded along a common extraction schema covering methodological family, learning paradigm, target deployment environment, evaluation protocol and the limitation its own authors acknowledge. Six families are identified and characterised: classical and ensemble pipelines, deep and hybrid architectures, outlier and semi-supervised formulations, reinforcement-learning agents, representation-learning front-ends, and an emerging group combining large language models, neuro-symbolic reasoning, temporal graphs, moving-target defence and federated learning. The profile of the corpus shows a decisive shift, with the emerging family absent before 2025 and accounting for nine of the seventeen studies dated 2026. The synthesis identifies five recurring weaknesses in evaluation practice: unverified corpus separability, inconsistent leakage control, near-absent calibration and cost reporting, single-run experimentation without significance testing, and the almost complete absence of cross-corpus transfer evidence. These are formulated as an agenda of open problems rather than as a ranking of methods, since the reported figures across the corpus are not commensurable.





