Deep Reinforcement Learning For SLA-Aware Cloud Resource Management in Financial Trading Platforms
Keywords:
Deep Reinforcement Learning, Cloud Resource Management, Service-Level Agreement, Financial Services, Autonomic Computing, Retrieval-Augmented Generation, Graph Neural Networks, Multi-Agent Reinforcement Learning, Explainable AI, Operational Resilience, Digital Operational Resilience ActAbstract
Cloud resource management for financial trading and portfolio-management platforms is dominated in production by a pattern that autoscaling literature rarely acknowledges: deliberate static over-provisioning as a trust posture, chosen because reactive autoscaling is judged too slow to react to the very market conditions during which the platform most needs to be available. This paper draws on production observations from two anonymized institutional-financial platforms, one a large US institutional asset manager serving over 50,000 users and more than one billion dollars in managed investments, the other a US fixed-assets management platform provider running ASP.NET WebAPI with SignalR real-time state on AWS EC2. Neither platform used any form of autoscaling; both accepted permanent cost inefficiency in exchange for guaranteed peak-hour capacity. Deep reinforcement learning offers a plausible path beyond this trade-off, but its deployment in regulated financial services faces two under-treated problems. The first is coordination across multi-application platforms sharing the same underlying infrastructure, where per-application independent scaling policies produce contention that a coordinated policy could pre-empt. The second is the auditability of automated decisions under the operational-resilience frameworks now in force, the EU Digital Operational Resilience Act (Regulation 2022/2554), the US SEC Regulation SCI (17 CFR § 242.1000-1007), and FINRA Rule 4370. This paper contributes a practitioner-grounded requirements analysis, a proposed DRL architecture that treats multi-application coordination as a first-class design property, a retrieval-augmented context substrate that lets the policy condition on top-k historical scenarios rather than on raw state alone, a graph-neural-network encoder over the application-dependency graph that produces a coordination-aware state representation, and an explicit discussion of the explainability barrier the paper partially addresses through case-based-reasoning audit trails derived from the retrieval substrate. A reference Python implementation and simulated evaluation accompany the architecture. Empirical evaluation on production workloads is future work.





