Contents
Figure 1: Two layers, two tools — the durable platform you review carefully, and the GPU capacity that changes hourly
Provisioning a GPU cluster by hand, clicking through a cloud console, is how you end up with an orphaned instance burning money three weeks after the experiment that needed it ended. ML teams specifically need two different categories of infrastructure tooling: general-purpose IaC for the surrounding platform — networking, storage, IAM, Kubernetes clusters — and something purpose-built for the actual GPU provisioning problem, which behaves nothing like provisioning a web server. Let's cover both.
A Quick Correction Before We Start
If you've seen CDK for Terraform (CDKTF) recommended anywhere, know that it was deprecated in December 2025 and its GitHub repository has since been archived. It let you write Terraform infrastructure in TypeScript, Python, Java, C#, or Go, transpiling your code into Terraform JSON before the Terraform CLI actually deployed it. That's no longer a live option for new projects — if you're currently running it, Pulumi offers tooling specifically for migrating CDKTF code and importing existing Terraform state, worth checking out before that project's lack of maintenance becomes an actual liability.
Terraform: Still the Default, With a Licensing Asterisk
Terraform remains the most widely adopted general-purpose IaC tool, and for ML teams specifically its enormous provider ecosystem matters a lot — every major cloud, plus specialized providers for things like SageMaker domains, Vertex AI resources, and Kubernetes-native ML tooling, already exist and are well-maintained. It's the right default for teams with existing Terraform expertise, established workflows, or genuinely multi-cloud infrastructure needs where provider coverage matters more than anything else.
The thing worth knowing explicitly: Terraform shifted to the Business Source License (BUSL-1.1) a while back, not fully open source in the traditional sense. That single decision is exactly what spawned OpenTofu, a community-driven fork maintained under the Linux Foundation, offering Terraform-compatible workflows and syntax under genuine open governance. If license terms matter to your organization, or you want assurance against a vendor unilaterally changing licensing terms again, OpenTofu is a drop-in-compatible alternative worth evaluating specifically for that reason, not because Terraform itself has gotten worse.
Terraform's HCL is deliberately limited as a language — no real loops beyond basic iteration constructs, no arbitrary control flow — and that's actually a feature for a lot of teams, not a limitation. Infrastructure diffs stay boring, flat, and reviewable by someone who isn't a strong general-purpose programmer, which matters a lot when your platform engineer on call at 2am needs to understand exactly what a terraform plan is about to do without reverse-engineering a junior developer's clever abstraction.
Pulumi: Infrastructure in a Real Programming Language
Pulumi takes a fundamentally different approach: you write infrastructure in TypeScript, Python, Go, C#, Java, or several other genuinely general-purpose languages, and Pulumi's engine deploys it directly against your cloud provider's APIs, with no separate compile-to-template step in between. That distinction matters more than it sounds: AWS CDK and the now-deprecated CDKTF both transpile your code into another system's format (CloudFormation or Terraform JSON respectively) before anything actually deploys, adding a layer of indirection Pulumi simply doesn't have.
For ML teams specifically, this pays off when your infrastructure logic genuinely needs real programming constructs — generating near-identical GPU cluster configurations across a dozen experiment variants, or building infrastructure provisioning directly into an internal tool your data scientists use, rather than keeping infrastructure code in an entirely separate repository and language from your application code. Pulumi also supports referencing existing Terraform modules directly without rewriting them, which matters if you've already got a meaningful Terraform investment and want to adopt Pulumi incrementally rather than as a rip-and-replace.
The honest trade-off: Pulumi's expressiveness is a double-edged sword. Infrastructure code written in a Turing-complete language can develop exactly the same complexity problems application code does — hidden control flow, clever abstractions only one engineer understands, a stray for-loop that quietly provisions 400 unintended resources during a late-night debugging session. Pick Pulumi if your team wants the same tests, linters, and IDE tooling they already use for application code, but budget for the code review discipline that infrastructure-as-real-code demands, since the guardrails HCL provides by being deliberately limited don't exist here.
AWS CDK: The Right Call If You're All-In on AWS
If your ML infrastructure lives entirely on AWS — SageMaker domains, GPU EC2 instances, S3 feature stores — and you want typed, high-level constructs rather than hand-written CloudFormation templates, AWS CDK is a genuinely strong, AWS-maintained choice. It compiles to CloudFormation under the hood, so you inherit CloudFormation's native AWS integration and state management rather than managing a separate state backend yourself, at the cost of that same compile-then-deploy indirection that makes Pulumi's direct-deploy model appealing by comparison.
CDK makes less sense the moment multi-cloud enters the picture: it's explicitly AWS-only, so if your organization runs ML workloads across AWS and GCP, or is evaluating cheaper GPU capacity from a neocloud provider, CDK alone won't cover that ground — and portability discipline is exactly what the multi-cloud MLOps piece argues you should be preserving deliberately anyway.
The ML-Specific Gap: General IaC Doesn't Solve GPU Provisioning
Here's the thing none of the tools above handle well on their own: provisioning a GPU for training is a genuinely different problem than provisioning a web server. GPU availability is volatile, spot markets swing wildly, desirable instance types are frequently unavailable on any single cloud, and the cost difference between spot and on-demand pricing, or between providers entirely, routinely exceeds 3x. Writing Terraform or Pulumi code that handles "try AWS, fall back to GCP, fall back to a neocloud, and automatically recover if the spot instance gets preempted mid-training" is possible, but it's reinventing a wheel that already exists.
That's the specific gap SkyPilot fills. It's an open-source orchestration layer purpose-built for ML workloads specifically: you define a job in a YAML file — GPUs needed, memory, disk — and SkyPilot automatically selects the cheapest available cloud and region across more than 20 providers, including AWS, GCP, Azure, Kubernetes clusters, and dedicated GPU clouds like Lambda, RunPod, and CoreWeave.
resources:
accelerators: A100:8
use_spot: true
run: |
python train.py
sky launch train.yaml -n my-training-job --use-spot
That single command handles selecting the best available provider based on current pricing and capacity, provisioning the instance, syncing your codebase to it, and, critically, automatically resuming the job from checkpoint if a spot instance gets preempted mid-run. Spot GPU instances typically run 60 to 80% cheaper than on-demand pricing, but per Cast AI's 2026 report, fewer than 2% of GPU workloads currently actually run on spot, largely because handling preemption and recovery manually is enough of a headache that most teams just don't bother. SkyPilot exists specifically to make that savings potential practical rather than theoretical.
A Real Production Example
Shopify runs SkyPilot for essentially all of its ML training workloads, using it as a launcher on top of persistent Kubernetes clusters running across multiple clouds — Nebius for H200s with InfiniBand interconnect for serious distributed training, GCP for additional capacity. The appeal they cite is concrete: jobs request disk space directly in their YAML, storage appears, and volumes get automatically cleaned up after seven days of disuse — no provisioning tickets, no orphaned volumes quietly eating budget the way manually-provisioned infrastructure tends to accumulate over time.
That orphaned-volume detail is worth sitting with. It's exactly the kind of cost leak that cost monitoring dashboards, covered elsewhere in this series, are built to catch after the fact — SkyPilot's automatic cleanup prevents it from happening in the first place.
SkyPilot vs a Platform Like Cast AI
Worth distinguishing these, since they're sometimes compared head to head despite solving different problems. SkyPilot is fundamentally a job orchestrator: you submit a training job, it finds the cheapest GPU across clouds, runs it to completion, and tears down afterward. Cast AI, covered in the cost monitoring piece in this series, is built for long-running Kubernetes services — inference endpoints, vector databases — things that stay up continuously rather than running to completion. If your workload is training or batch jobs, SkyPilot fits naturally; if it's a persistently running inference service, a platform like Cast AI's autoscaling and bin-packing is the better-suited tool.
How These Tools Fit Together in Practice
For most ML teams, the realistic answer isn't choosing one tool, it's layering them by responsibility. Terraform, OpenTofu, or Pulumi provision your durable platform infrastructure: VPCs, IAM roles, persistent Kubernetes clusters, S3 buckets — the things that don't change job to job and genuinely benefit from careful state management and change review. SkyPilot sits on top, handling the actual GPU provisioning for training and batch jobs, where availability and pricing change hour to hour and a general-purpose IaC tool's deliberate, careful-review workflow is the wrong fit for something you want spun up and torn down in minutes, automatically, based on whatever capacity happens to be cheapest right now.
This mirrors the Terraform-for-ML-infrastructure pattern covered elsewhere in this series directly: use Terraform or Pulumi for the SageMaker domain or the persistent GPU cluster shell, and let a tool purpose-built for ephemeral GPU scheduling handle what actually runs inside it — the same separation of concerns the CI/CD rollback discipline applies at the deployment level.
Choosing Based on Your Actual Situation
If your team already knows Terraform and you're multi-cloud, stick with it, or move to OpenTofu specifically if the BUSL licensing terms are a genuine concern for your organization — migration is close to seamless given OpenTofu's deliberate compatibility.
If your engineers want infrastructure written in the same languages, with the same tests and IDE tooling, they already use for application code, and you're willing to invest in the code review discipline a Turing-complete infrastructure language demands, Pulumi is the stronger fit.
If you're exclusively on AWS and want typed, high-level constructs without introducing a third-party tool or ongoing licensing cost, AWS CDK covers that ground well, accepting its AWS-only scope and the CloudFormation compile step as the trade-off.
If GPU provisioning itself, not the surrounding platform, is your actual pain point — spot preemption handling, multi-cloud GPU price arbitrage, queueing hyperparameter sweeps across dozens of trials — add SkyPilot specifically for that layer, regardless of which general-purpose IaC tool handles everything underneath it.
Common Pitfalls
- Trying to handle GPU spot preemption and multi-cloud GPU arbitrage by hand-rolling it in Terraform or Pulumi is a real time sink — this is precisely the problem SkyPilot already solved well, don't reinvent it.
- Ignoring Terraform's BUSL licensing shift without checking whether it actually matters for your organization, or without evaluating OpenTofu as the direct response, means making an infrastructure decision without the full picture.
- Adopting Pulumi's full programming-language flexibility without adding the code review rigor that comes with it leads to exactly the "one engineer's clever abstraction nobody else understands" problem that HCL's deliberate limitations were designed to prevent.
- Still running CDKTF for new projects given its December 2025 deprecation and subsequent archival is building on infrastructure tooling nobody's maintaining anymore — migrate off it deliberately rather than discovering the gap during an incident.
Recommended Books
| Cover | Book | Description | Get it |
|---|---|---|---|
![]() |
Terraform: Up & Running | the practical reference for the layer most ML teams start with: state management, module design, and reviewable infrastructure change workflows. | View on Amazon |
![]() |
Infrastructure as Code | the principles behind the tools: treating servers, networks, and tooling as software, with patterns for immutable and ephemeral infrastructure that map directly onto the GPU layer discussed here. | View on Amazon |
![]() |
Kubernetes in Action | since the persistent cluster shell most of these tools provision is a Kubernetes cluster, and understanding what the platform layer actually manages makes the SkyPilot-on-top split obvious. | View on Amazon |
Unlock AI That Actually Works
Get lifetime access to GPT-6 Astra, Claude Fable 5.1, Gemini 3.5, Grok 4.5, and more — all in one platform. Build websites, apps, videos, content, and digital products from a single command. No monthly fees. No tool-hopping.
Click here to get GPTAstra Max now — one-time payment, lifetime access.
Frequently Asked Questions
What is the best IaC tool for an ML team in 2026?
Terraform remains the default for the durable platform layer — VPCs, IAM, Kubernetes clusters, storage — because of its provider ecosystem, with OpenTofu as the license-clean drop-in if the BUSL terms matter to your organization, and Pulumi the stronger fit when your team wants infrastructure in the same languages and with the same tests as application code. GPU provisioning for training jobs is a separate layer that general-purpose IaC handles badly; that's where SkyPilot comes in.
What happened to CDK for Terraform (CDKTF)?
CDKTF was deprecated in December 2025 and its GitHub repository has since been archived. It let you write Terraform infrastructure in TypeScript, Python, Java, C#, or Go and transpile it into Terraform JSON before the CLI deployed it, but it is no longer a live option for new projects. Pulumi offers tooling for migrating CDKTF code and importing existing Terraform state if you're still running it.
Why isn't Terraform enough for GPU provisioning?
GPU availability is volatile, spot markets swing wildly, desirable instance types are frequently unavailable on any single cloud, and the price gap between spot and on-demand — or between providers — routinely exceeds 3x. Hand-writing "try AWS, fall back to GCP, fall back to a neocloud, recover from preemption mid-training" in HCL reinvents an existing wheel: SkyPilot does the selection, provisioning, code sync, and checkpoint resume in one command across more than 20 providers.
Should we use SkyPilot or Cast AI for our workloads?
They solve different problems. SkyPilot is a job orchestrator: submit a training or batch job, it finds the cheapest available GPU across clouds, runs it to completion, and tears it down. Cast AI targets long-running Kubernetes services — inference endpoints, vector databases — with autoscaling and bin-packing for things that stay up continuously. Training and batch work fits SkyPilot; a persistently running inference service fits a Kubernetes cost platform.
Wrapping Up
ML teams need two layers of infrastructure tooling working together: Terraform, OpenTofu, Pulumi, or AWS CDK for the durable platform underneath, and SkyPilot specifically for the volatile, cost-sensitive GPU provisioning layer on top, where availability changes hourly and automatic spot recovery genuinely matters. Pick your general-purpose tool based on team language preference and cloud footprint, and add SkyPilot regardless of that choice the moment GPU cost or availability becomes a real operational pain point.
Will one tool cover everything cleanly? No, and trying to force general-purpose IaC to handle ML-specific GPU scheduling, or trying to make SkyPilot provision your VPCs and IAM policies, both miss what each tool is actually built for. But layer them by responsibility — durable platform in one, ephemeral GPU jobs in the other — and you get both the careful, reviewable infrastructure changes your platform needs and the fast, cost-aware GPU provisioning your training jobs actually want.


