On October 1, 2026, Modal announced on its official engineering blog that Modal Clusters, its multi-node GPU cluster product, is now generally available (GA). Developers only need to add a single decorator, modal.clustered, to a Python function to get a multi-node cluster: nodes are interconnected via InfiniBand verbs at up to 6.4 Tbps, with PyTorch and NCCL configured automatically. A cluster is ready within seconds and billed by the second.
What you get beyond one decorator
modal.clustered is not a standalone piece of syntactic sugar. Modal Clusters integrates with all of Modal's existing primitives: training checkpoints go to Volumes, data is mounted in through Cloud Bucket Mounts, and jobs are orchestrated with Queues. The trickiest part of the networking layer is folded into a single flag, rdma=True — drivers, environment variables, and userspace libraries, which differ across cloud providers, are handled automatically by the platform, so PyTorch + NCCL works out of the box.
On pricing, traditional multi-node setups either charge by the hour or require capacity reservations, and you pay even when the machines sit idle. Modal is sticking to its serverless playbook here: run whenever you want, and pay only for what you actually use.
Background: where 1.5 years of engineering went
According to Modal, the past 1.5 years were spent battle-testing this capability, with the core work being a rebuild of the scheduler and the networking stack.
Traditional schedulers are greedy: each machine polls the task queue on its own, and the scheduler checks for fitting work one item at a time. But multi-node scheduling changes the atomic unit from "one machine" to "N machines at once," so Modal wrote a new gang scheduler: it first surveys the entire fleet and every pending cluster, makes one global placement plan, and then acts on it, grouping nodes in the same availability zone onto the same RDMA fabric.
RDMA itself is notoriously difficult to support. Modal opted for gVisor, the more secure multi-tenant container runtime, which had no RDMA support — so Modal built it in itself and contributed the changes back to the open-source community.
Who is already using it
The blog names three workloads that ran on the platform during the testing period: customer-experience agent company Decagon fine-tuned open models at the trillion-parameter scale; humanoid robotics company 1x pre-trains NEO's world model on multi-node B300 clusters, spinning up hundreds of GPUs for large experiment sweeps; and Runway runs multi-node inference for its latest video generation and editing models, Gen-4.5 and Aleph 2.0, with the autoscaler scaling those clusters as a unit.
Who should adopt it, and who can wait
Companies with multi-node training or inference needs that don't want to build and operate their own cluster team are the most direct audience: post-training, trillion-parameter fine-tuning, and cross-node inference splitting previously meant either staffing an infra team or living with cloud providers' queues and reservations.
If your workloads always fit on a single GPU, Modal Clusters is irrelevant to you; if you already have long-term reserved on-prem capacity, migrating may not pay off either. Modal Clusters opens up an option that barely existed before — renting multi-node capacity by the second — and its value first belongs to teams whose compute demand spikes and dips.