Performance Comparison of J48 and Random Forest Algorithms for Network Intrusion Detection Using the KDD Cup 1999 Dataset

Authors

  • Pratik Jain
  • Archana Liladhar Rane
  • Madhav Mohan Vagmare
  • Rajni Jainwal
  • Dr. Nita Shinde
  • Javed Rauf Attar

Keywords:

Intrusion Detection System (IDS), Network Security, Machine Learning, Supervised Learning, Decision Tree, J48 Algorithm, Random Forest, WEKA, KDD Cup 1999, Cybersecurity, Classification, Data Mining.

Abstract

The proliferation of computer networks, cloud services, Internet services, and electronic communications has led to an explosion of cyber threats in quantity and complexity. Attacks like Denial-of-Service (DoS), Probe, Remote-to-Local (R2L), User-to-Root (U2R) are attacking the well-organized enterprises worldwide day by day, and these traditional security techniques like firewall, antivirus have limitations to sniff out the well sophisticated or zero-day attacks. As a result, Intrusion Detection Systems (IDSs) play a key role in contemporary cybersecurity frameworks by monitoring network traffic in real-time and detecting malicious activities.

Due to its ability to learn patterns from historical network traffic and classify normal and malicious network connections with high accuracy, machine learning has become a promising tool for enhancing intrusion detection. Amongst supervised learning algorithms, Decision Tree (J48) and Random Forest are some of the popular ones that are extensively applied based on its classifying capabilities, computational complexity and ease in utilizing the algorithm. Despite  J48 producing a single decision tree and thus interpretable decision rules, Random Forest utilizes ensemble learning to average many decision trees to obtain more accurate predictions and less overfitting.

In this work, we compare the performances of classification algorithms J48 and Random Forest in detection of network intrusions based on the KDD Cup 1999 dataset. The experiments are performed in the data mining environment of WEKA 3. 8 under the same experimental setup to assure the compatibility of the study. The dataset is preprocessed to remove outliers and missing values before classification, and to enable the data set to be used as supervised learning materials. Both algorithms are tested against the conventional performance measure like classification accuracy, precision, recall, F-measure, ROC Area, confusion matrix, model building time and computational cost. In this work, we evaluate different classifiers and their combinations aiming at the best detection performance with a feasible training and testing time (computational feasibility), with the overall goal to recommend the classifier with best performance for practical intrusion detection systems. This comparative study enables a better understanding of the merits and demerits of both the algorithms, and also helps the researchers and cyber security experts for choosing a suitable supervised machine learning algorithm based on the required parameters for network intrusion detection.

Downloads

Published

2026-09-05

How to Cite

Jain, P., Rane, A. L., Vagmare, M. M., Jainwal, R., Shinde, D. N., & Attar, J. R. (2026). Performance Comparison of J48 and Random Forest Algorithms for Network Intrusion Detection Using the KDD Cup 1999 Dataset. International Journal of Artificial Intelligence and Machine Learning, 6(9s), 1130–1142. Retrieved from https://www.svedbergopen.com/index.php/ijaiml/article/view/1573