Intelligent Defect Classification and Root-Cause Prediction for Quality Assurance in Regulated Enterprise and Healthcare Systems

Authors

  • Pratik Dinkar Rane

Keywords:

Defect Classification, Root-Cause Analysis, Explainable AI, Software Quality Assurance, Healthcare QA, HIPAA Compliance, Machine Learning, Enterprise Testing

Abstract

Manual defect triage in enterprise and healthcare software systems consumes an estimated 20 to 30 percent of QA engineering time. Triage activities include reading failure output, classifying defect type, identifying the responsible subsystem, estimating priority, and routing the defect to the correct engineering team. In high-volume regulated environments, this overhead constrains sprint velocity and creates a bottleneck that delays defect remediation. Existing defect classification research primarily targets pre-release defect prediction and binary categorization of post-release defects, leaving the post-failure automated triage problem largely unaddressed for regulated enterprise and healthcare contexts. This paper introduces PRAXIS, a four-layer post-failure intelligence framework for automated defect classification and root-cause prediction. PRAXIS has a seven-category compliance-aware taxonomy, a classification pipeline for natural language processing with four deployment modes (cloud LLM and local traditional ML options), a three-engine ensemble root-cause prediction model, and a Triage and Compliance Reporting Layer that links classification outcomes to relevant regulatory control identifiers across HIPAA, FHIR, 21 CFR Part 11, and SOX frameworks. The framework accepts test failure signals from JUnit XML, SARIF, APM/OTLP, git log, and defect tracker connectors without requiring modification of existing test infrastructure. An exploratory evaluation on a 200-signal synthetic dataset covering SAP BTP, Salesforce CRM, and FHIR healthcare domains shows 84% weighted-average classification accuracy, 82% accuracy in routing subsystem ownership, and a mean triage time reduction from about 15 minutes to under 30 seconds. An ensemble ablation study confirms that the three-engine ensemble achieves 72% top-3 root-cause accuracy, outperforming all individual engines and engine pairs. Expert review of 35 output samples by senior QA engineers yields mean plausibility ratings of 3.9 to 4.3 on a 5-point Likert scale with inter-rater agreement of Krippendorff alpha = 0.74.

Downloads

Published

2026-09-05

How to Cite

Rane, P. D. (2026). Intelligent Defect Classification and Root-Cause Prediction for Quality Assurance in Regulated Enterprise and Healthcare Systems. International Journal of Artificial Intelligence and Machine Learning, 6(9s), 795–809. Retrieved from https://www.svedbergopen.com/index.php/ijaiml/article/view/1544