Consistency-Aware and Secure Autoscaling of Stateful Microservices under Replication Lag and Failure-Recovery Constraints
Keywords:
Autoscaling; Stateful microservices; MongoDB; Replication lag; Consistency; Failure recovery; Kubernetes; OAuth 2.0; Java security; Spring SecurityAbstract
Horizontal scaler abstractions assume all instances of microservice replicas are identical resources․ These abstractions do not apply to stateful microservices using MongoDB replica sets․ Additional MongoDB secondary replicas can only be used after a long initial sync process that burdens other replica set members․ Multiple read-only secondary replicas become stale compared to the primary and return outdated information․ Primary failures prevent write access during elections and topology changes․ This paper introduces CAAS‚ a Consistency-Aware AutoScaler for stateful microservices using MongoDB․ CAAS implements the MAPE-K control loop with four new components․ A lag detection component measures available read capacity only from secondaries that are not stale․ A sync planning component predicts future demand across the full sync process duration and allows only one sync process at a time‚ starting with the least busy replica set member․ An apply-headroom component reserves capacity to handle secondary oplog resynchronization․ A recovery prevention component prevents unsafe scaling down after a failure․ CAAS is secured with OAuth 2.0 and Java security: a Spring Security resource server validates JWT access tokens, token-revocation and session reads are served only from majority-committed data, and members authenticate with X.509 certificates over TLS. The design also applies to other leader–follower stores such as PostgreSQL and MySQL. Simulation experiments with a Python discrete-event simulator and three YCSB-inspired workloads with injected failures across ten random seeds show that CAAS reduces read SLO violations in a flash-crowd workload by 99․64 percentage points‚ from 11․41% to 0․36% of time‚ compared to other approaches․ Stale reads are reduced by 99․44 percentage points‚ from 5․68% to 0․03%‚ and recovery time is reduced by 82․5 seconds‚ from 87․8 to 15․3 seconds․ Read-heavy workloads using CAAS require 7% to 17% more nodes․ Write-heavy workloads benefit more from a stale-read-free HPA approach for reads and CAAS for writes․ Considerations related to MongoDB replication lag and recovery state are important for autoscaling stateful microservices․





