Contents
Figure 1: Portability isn't a product you buy — it's a set of boundaries you draw deliberately, one layer at a time
Recall the Snowflake vs BigQuery vs Redshift article's genuinely sobering number from earlier in this series: migrating a 50TB warehouse has been benchmarked at $80K-$250K in engineering time, and egress fees alone can run $90-150 per terabyte moved out. A 2026 cost-optimization analysis puts a number on the broader pattern that figure was one instance of: vendor lock-in inflates multi-cloud costs by 20-30% overall, through proprietary APIs, egress fees, and retraining overhead combined. This article is about the architectural discipline that keeps that number from compounding specifically in your MLOps stack — not cloud strategy in the abstract, but the concrete pieces this series' MLOps arc already built.
Here's the nuance worth getting right immediately, because most guides on this topic miss it: Kubernetes is described directly in current analysis as "both the solution and a potential source of lock-in." Recall the MLOps platforms article's Kubeflow coverage from earlier in this series directly — it's named explicitly in current research as instantiating exactly this portability strategy for MLOps specifically, through standardized APIs and modular, selectively-adoptable components. But a Kubernetes cluster riddled with EKS-specific or AKS-specific managed add-ons is genuinely no more portable than a fully proprietary stack — the orchestrator alone doesn't guarantee anything.
By the end of this guide, you'll understand the four-layer portability stack current practice converges on, which specific pieces of this series' MLOps content already support it and which don't, and the pragmatic exceptions worth making deliberately rather than by accident. IMO, "keep peripheral services native to their cloud where appropriate" is genuinely the single most important caveat to the whole multi-cloud-purist instinct :)
The Four-Layer Portability Stack
Current analysis converges consistently on the same four layers, worth treating as a checklist rather than a vague principle: Kubernetes handles compute portability, open databases handle data portability, object-storage abstractions handle storage portability, and Infrastructure-as-Code (Terraform) handles deployment portability. None of these four alone is sufficient — recall this being the exact same "specialized tools, each doing one job" theme running through this series' entire MLOps arc.
Layer One: Kubernetes for Compute — With the Caveat Stated Directly
Recall the MLOps platforms article's Kubernetes-native stack coverage directly — Kubeflow, KServe, Kueue, KEDA, all explicitly named there as the compute orchestration layer for MLOps specifically, and all of it catalogued in the MLOps platforms comparison from earlier in this series.
- Kubernetes abstracts infrastructure via standard APIs, allowing workload portability across AWS EKS, Azure AKS, and on-premise clusters without rewriting application deployment configurations — genuinely the mechanism making "train on one cloud, serve on another" a realistic option rather than a rewrite.
- The explicit caveat worth repeating: Kubernetes deserves special attention specifically because it's both the solution and a potential lock-in source. A cluster built around a cloud provider's managed add-ons — AWS's specific load balancer controller, Azure's specific identity integration — genuinely recreates the same lock-in problem one layer up the stack.
- A concrete portability target worth adopting: track the percentage of compute genuinely running on standard Kubernetes versus provider-specific managed services, with 75%+ on Kubernetes named directly as a reasonable quarterly KPI target in current guidance.
Layer Two: Open Data Formats — The Highest-Leverage Layer
Recall the data lake versus data warehouse argument from earlier in this series: open table formats (Iceberg, Delta Lake, Hudi) are explicitly what let different query engines access the same underlying data without full migration. This is genuinely the same mechanism, now framed specifically as a vendor-lock-in defense rather than just a lakehouse architecture choice.
- Store data in formats that aren't proprietary — Parquet, Avro, Delta, Iceberg — rather than a warehouse's own internal storage format. Recall the Snowflake vs BigQuery vs Redshift warehouse comparison's egress and SQL-dialect-lock-in warning directly — this is the concrete defense against exactly that cost.
- A genuinely important, specific recommendation from current research: insist on open-format tables and an exportable catalog regardless of which platform you choose today — this is precisely the same advice that warehouse comparison article gave, now generalized as a cross-cloud MLOps principle rather than a single-vendor decision.
- DVC's own design philosophy already supports this — pointer files in Git, actual data payload in whatever S3-compatible remote you configure; swapping that remote between providers is genuinely a configuration change, not a data migration.
Layer Three: Object-Storage Abstraction
S3-compatible storage protocols have become genuinely universal enough that MinIO, Azure Blob (via compatibility layers), and GCS can often be addressed through the same client code — worth treating storage-layer portability as a distinct concern from the data-format question above, since you can have open Parquet files sitting in a genuinely proprietary, hard-to-migrate storage API. The format layer and the storage layer fail independently, which is exactly why current checklists list them as separate layers rather than one.
# Same client config works against S3, MinIO, and every S3-compatible backend
storage:
endpoint: ${STORAGE_ENDPOINT} # https://s3.amazonaws.com | http://minio:9000
region: ${STORAGE_REGION}
bucket: ml-artifacts
style: path # path-style addressing keeps portability
Layer Four: Infrastructure as Code — With Its Own Caveat
Terraform manages multi-cloud resources, but uses provider-specific modules — worth stating this limitation directly rather than assuming "we use Terraform" automatically means "we're portable." A Terraform configuration built entirely from AWS-specific resource blocks is genuinely no more portable than hand-clicking through the AWS console; the portability comes specifically from writing provider-agnostic modules or maintaining parallel provider-specific modules behind a consistent interface, not from the tool choice alone.
Where This Series' MLOps Arc Already Supports Multi-Cloud, and Where It Doesn't
Worth auditing this explicitly against the specific tools this series already covered, rather than treating multi-cloud strategy as an entirely separate topic from everything built up across the MLOps arc.
- MLflow (recall the model registry article directly) — genuinely cloud-agnostic by design; its tracking server, artifact store, and registry can all point at any S3-compatible backend regardless of provider, making it a naturally portable choice for exactly this reason.
- The SageMaker/Vertex AI/Azure ML integrated platforms (recall the MLOps platforms comparison article directly) — these are explicitly the opposite end of this spectrum; recall that article's own honest framing: "the tradeoff is real lock-in — migrating off SageMaker later means genuinely rebuilding pipeline logic against a different platform's abstractions, not just changing a config file." Choosing one of these is a deliberate lock-in tradeoff for integration convenience, not an accident — worth making that choice knowingly rather than by default.
- BentoML and KServe for serving (recall that same article directly) — genuinely portable serving layers, since both run on standard Kubernetes rather than a cloud-specific managed inference service, directly supporting the "custom prediction APIs can route requests to multiple backend implementations, enabling gradual migration or multi-vendor deployments" pattern current research recommends.
- Kubeflow's modular architecture (recall this being named explicitly in current multi-cloud-lock-in research) — delivers "Kubernetes-native MLOps with modular component architecture, enabling selective adoption without platform lock-in," genuinely the direct MLOps-specific validation of the broader Kubernetes-abstraction strategy.
Two Genuine Architectural Patterns, Not One
Current practice clusters around two distinct patterns, and conflating them leads to the wrong implementation effort.
Workload placement — each application or pipeline stage lives entirely in whichever cloud fits it best (train on GCP for TPU access, serve on AWS where your application layer already runs), with genuine data movement and integration work between the two, but no attempt at making any single workload itself cloud-agnostic.
Multi-cloud Kubernetes — a consistent platform runs identically across providers, with workloads placed by policy rather than by hard-coded cloud assignment, offering genuinely maximum portability at the cost of real, sustained platform engineering investment.
The honest, pragmatic guidance worth adopting directly: use managed services where they genuinely accelerate delivery, and maintain portability specifically for the parts of the system most likely to actually need it — keep peripheral services native to their cloud where appropriate, rather than paying the multi-cloud-Kubernetes tax for every single component regardless of its actual migration risk.
Observability: OpenTelemetry as the Fourth Pillar Beyond the Core Four
Worth naming directly as a genuinely distinct concern from compute, data, storage, and deployment: set up OpenTelemetry specifically so logs, metrics, and traces carry no vendor lock-in — recall the model monitoring article's Evidently AI and drift-detection coverage directly; the monitoring logic in that article is already tool-agnostic, but the underlying telemetry pipeline feeding it needs the same portability discipline as everything else in this stack, or your observability data itself becomes a migration cost nobody budgeted for.
The practical shape of that: one collector, OTLP on the wire, backends chosen afterwards and swappable without touching instrumented code.
receivers:
otlp:
protocols: { grpc: null, http: null }
exporters:
otlphttp:
endpoint: ${BACKEND_OTLP_URL} # point at any OTLP-speaking backend
processors: [batch, memory_limiter]
service:
pipelines:
metrics: { receivers: [otlp], exporters: [otlphttp] }
traces: { receivers: [otlp], exporters: [otlphttp] }
Instrument once with the OpenTelemetry SDK, and the choice of where telemetry lands becomes an infrastructure decision rather than a rewrite — the same separation the model monitoring piece assumes when it swaps drift-detection tooling underneath a fixed metric contract.
The Regulatory Layer: The Compliance Thread Runs Through Here Too
Worth connecting directly to the EU AI Act mentions from the model monitoring, federated learning, and synthetic data privacy articles earlier in this series — current guidance specifically names being "ready for the EU Data Act portability mandate" as a concrete, near-term driver for this entire multi-cloud discipline, not just a cost-optimization nicety.
The same regulatory pressure pushing teams toward federated learning and differentially private synthetic data is genuinely pushing infrastructure architecture toward data portability requirements too — this is one connected compliance story, not three separate ones. Data residency rules, subject-access exports, and model provenance records all get dramatically cheaper when your data layer already lives in open formats behind an exportable catalog.
A Practical Audit Sequence
- Audit your current architecture first — identify exactly which vendor-specific dependencies exist today, the universally-recommended first step across every current source on this topic, before designing anything new.
- Classify workloads by actual migration risk, not by default assumption — recall the workload-placement-versus-multi-cloud-Kubernetes distinction directly; not every component needs maximum portability investment.
- Check your data layer specifically against the open-format standard — this is genuinely the highest-leverage layer to get right, since data is explicitly named as "often the hardest part to move."
- Confirm your IaC modules are genuinely provider-agnostic, not just "written in Terraform" — the tool alone doesn't guarantee portability; the module design does, and the fastest way to see the truth is counting provider-prefixed resource blocks:
# How provider-specific is this Terraform tree, really?
grep -rhoE 'resource "aws_[a-z0-9_]+"' infra/ | wc -l
grep -rhoE 'resource "google_[a-z0-9_]+"' infra/ | wc -l
grep -rhoE 'resource "azurerm_[a-z0-9_]+"' infra/ | wc -l
- Document and test an actual exit strategy from day one — recall the CI/CD article's rollback-planning discipline directly, applied here at the infrastructure-provider level rather than the model-deployment level; a "quarterly two-cloud deploy day" is explicitly named in current guidance as a concrete practice for keeping this capability genuinely tested rather than theoretical.
Common Mistakes People Make
- Assuming "we use Kubernetes" automatically means "we're not locked in." Recall the explicit caveat directly — a cluster built around provider-specific managed add-ons recreates the same lock-in one layer up.
- Treating Terraform adoption as sufficient without checking module portability. Recall the direct limitation — provider-specific Terraform modules are genuinely no more portable than the proprietary console workflows they replace.
- Choosing an integrated MLOps platform (SageMaker, Vertex AI) without acknowledging the deliberate lock-in tradeoff being made. Recall the MLOps platforms article's own honest framing directly — this is a legitimate choice for integration convenience, but it should be made knowingly, not by default.
- Pursuing multi-cloud-Kubernetes-everywhere for components with genuinely low migration risk. Recall the pragmatic "keep peripheral services native where appropriate" guidance directly — full portability investment isn't free, and not every piece needs it.
- Treating data portability as solved by Kubernetes alone. Recall the explicit layer separation directly — "Kubernetes handles compute portability (but not data)"; open formats and storage abstraction are genuinely separate, required layers.
Recommended Books
| Cover | Book | Description | Get it |
|---|---|---|---|
![]() |
Kubernetes in Action | the definitive guide to the compute layer this stack rests on, and the best way to understand exactly which parts of a cluster are standard and which are provider sugar. | View on Amazon |
![]() |
Infrastructure as Code | the deployment-portability layer in book form: managing servers, networks, and tooling through code rather than console clicks, with the module discipline the Terraform caveat demands. | View on Amazon |
![]() |
Terraform: Up & Running | the practical companion for writing the provider-agnostic modules that make "we use Terraform" mean something, from workspace layout to reusable module design. | View on Amazon |
Unlock AI That Actually Works
Get lifetime access to GPT-6 Astra, Claude Fable 5.1, Gemini 3.5, Grok 4.5, and more — all in one platform. Build websites, apps, videos, content, and digital products from a single command. No monthly fees. No tool-hopping.
Click here to get GPTAstra Max now — one-time payment, lifetime access.
Frequently Asked Questions
What are the four layers of a portable MLOps stack?
Kubernetes for compute portability, open data formats for data portability, object-storage abstraction for storage portability, and provider-agnostic Infrastructure as Code for deployment portability. None of the four is sufficient alone: Kubernetes moves your workloads but not your data, open Parquet files can still sit in a proprietary storage API, and a Terraform tree full of aws-specific resource blocks changes nothing about how hard a migration would be.
Does using Kubernetes guarantee multi-cloud portability?
No. Current analysis describes Kubernetes as both the solution and a potential source of lock-in: a cluster wired to provider-specific managed add-ons like AWS Load Balancer Controller or Azure's identity integration recreates the same lock-in one layer up the stack. Portability comes from standard APIs and upstream components, which is why guidance suggests tracking the share of compute running on standard Kubernetes rather than provider services, with 75%+ named as a reasonable quarterly target.
How much does vendor lock-in actually cost?
A 2026 cost-optimization analysis puts vendor lock-in at a 20-30% inflation of overall multi-cloud spend, driven by proprietary APIs, egress fees, and retraining overhead combined. Egress alone commonly runs $90-150 per terabyte moved out, and migrating a 50TB warehouse has been benchmarked at $80K-$250K in engineering time — the same compounding pattern, scaled to your own stack.
Should every workload in our stack be cloud-agnostic?
No, and treating it that way is one of the most common mistakes. The pragmatic guidance is to use managed services where they genuinely accelerate delivery and to invest in portability specifically where migration risk is highest — workload placement across clouds for the pieces that need it, and keep peripheral services native to their cloud where appropriate rather than paying the multi-cloud Kubernetes tax everywhere.
Wrapping Up
Multi-cloud MLOps portability genuinely rests on four distinct layers — Kubernetes for compute (with the explicit caveat that it can itself become a lock-in vector if built around provider-specific add-ons), open data formats for data portability, object-storage abstraction, and provider-agnostic Infrastructure as Code — and the tools this series' MLOps arc already covered split cleanly along exactly this line: MLflow, Kubeflow, BentoML, and KServe genuinely support it; SageMaker, Vertex AI, and Azure ML represent a deliberate, knowing trade of lock-in for integration convenience.
Remember that vendor lock-in's real cost — 20-30% inflated multi-cloud spend through proprietary APIs, egress fees, and retraining overhead — compounds exactly the same way the $80K-250K warehouse migration figure from earlier in this series did, and that the pragmatic answer isn't maximum portability everywhere, it's deliberate portability investment specifically where migration risk is genuinely highest. FYI, this article genuinely closes a thread that's run quietly through this entire series — the warehouse egress fees, the open table format discussion, the Kubernetes-native MLOps platforms, and the EU regulatory pressure toward data portability were each one piece of this same story, told from a different angle each time :)
Now go run the audit this article's first step recommends — check whether your own project's current architecture, built from any of the tools covered across this series' MLOps arc, leans toward MLflow-and-Kubeflow's inherent portability or SageMaker-and-Vertex's deliberate convenience tradeoff. That answer, more than any abstract multi-cloud strategy, is what should drive your next infrastructure decision.


