Absorbing Generative AI into an Existing ML Platform: An Integration Pattern for a Shared Control Plane
Keywords:
Control Plane; Generative AI Serving; Large Language Model Integration; Machine Learning Platform Architecture; Model Governance.Abstract
Meeting generative-AI demand by standing up a separate stack beside an existing machine learning platform is a common industry response, but it tends to duplicate governance, observability, registry, lineage, rollout, and cost workflows. This paper reports an architectural pattern that instead extends the existing platform to serve both predictive and generative workloads, through a shared control plane paired with runtime-specific serving lanes. The design is organized as a 2x3 matrix: two model families (classical, generative) crossed with three serving lanes (online, near-real-time, batch). Wire-level contracts shared across all six matrix cells are identified, along with the generative-AI-specific extensions needed for serving, retrieval-augmented generation, and parameter-efficient fine-tuning, and the operational changes required for evaluation, safety, and cost attribution. The contribution claimed is an integration pattern and design recipe rather than a new inference engine, scheduler, or serving algorithm; the underlying serving mechanics, including paged attention, disaggregated prefill and decode, and low-rank adaptation, are adopted from established literature and are not claimed as novel here. A Kubernetes-based reference architecture is described, the small set of net-new control-plane code is separated from the substrate it reuses unmodified, and the evidence basis, trade-offs, and limitations are discussed for platform teams considering a similar transition. Every quantitative claim is labeled by evidence category (design target, illustrative sizing, or extrapolation) since the design has not yet been benchmarked against a production workload.





