An Autonomous Decision-Making Framework For GPU Fabric Fault Remediation: Intelligent Systems For AI Training Infrastructure

Authors

  • Indra Kumar Mondal

Keywords:

Autonomous Infrastructure Operations, GPU Fabric Management, Decision-Making Framework, AIOps, Human-in-the-Loop Systems, Intelligent Fault Remediation.

Abstract

Large-scale AI training depends on GPU fabrics that behave less like conventional networks and more like a single distributed processor, where a fault in one link can stall an entire synchronized job. As training clusters scale into the tens of thousands of GPUs, mean time between failures falls sharply one recent study found mean-time-to-failure declining by roughly two orders of magnitude between an eight-GPU job and a 1,024-GPU job [1] and the old model of a human operator watching dashboards no longer matches the speed or volume of faults a modern fabric produces. This article proposes a tiered, bounded-autonomy framework for GPU fabric fault remediation, in which autonomous action is scoped not by fault type but by the reversibility and blast radius of the corrective action. The framework organizes remediation into four tiers: detection, diagnosis, recommended action, and autonomous action, with escalation to a human operator triggered whenever an action is large, hard to reverse, or matches a failure signature the system has not previously validated. Drawing on production evidence from large-scale GPU cluster operators and established human-automation interaction models, the article argues this reversibility-gated boundary fits GPU fabric operations better than a generic autonomy maturity ladder, since failure costs here are measured directly in GPU-hours. The discussion compares this framework against existing fault-detection systems, which already report strong diagnostic accuracy in production RoCE fabrics [4, 12], and against human-AI collaboration research in other safety-critical settings, where time-bounded human confirmation of AI recommendations has been shown to improve outcomes [2]. The aim is to guide how the next generation of self-operating AI training infrastructure expands autonomy deliberately, as trust in a defined, auditable authority boundary is earned.

Downloads

Published

2026-09-05

How to Cite

Mondal , I. K. (2026). An Autonomous Decision-Making Framework For GPU Fabric Fault Remediation: Intelligent Systems For AI Training Infrastructure. International Journal of Artificial Intelligence and Machine Learning, 6(9s), 1487–1499. Retrieved from https://www.svedbergopen.com/index.php/ijaiml/article/view/1605