Blog

Towards Fluid Compute - Making Python Sessions Portable

The machine is not the computation. It is just where the computation happens to be running right now.

Engineering13 min readEldar Hasanov

Fluid compute is the idea that a workload and its execution environment are decoupled, meaning the former no longer depends on the latter. In a utopian fluid compute scenario, your program would run on arbitrary hardware that scales and changes based on the workload's requirements.

To some extent, cloud environments and managed on-premises clusters are implementing early versions of this utopia, but I have been exploring how to push the idea further as part of @clusyio, starting with Python notebooks.

Python sessions have an important design property that we subconsciously accept without much thought.

You start a process, import some libraries, load a dataset, build a few DataFrames, fit a model, download some files, maybe train a model for an hour or two. Slowly but surely, the process becomes valuable.

Importantly, not just the source code, but the running process itself.

It now contains a computational state that may have taken hours to construct. Specifically, there are objects in memory, models and optimizers, random-number generators, cached results, imported packages, files created along the way, and all the relationships between them.

And all of that state belongs to one machine and execution environment you started with.

Now imagine if that machine disappears or breaks, gets preempted (in a cluster scenario), runs out of GPU memory, or simply stops being the right machine for the next part of the workload. In those cases, we either move and rerun from the beginning or generally fall back to some version of the same workflow: save whatever we can, provision something else, reconstruct the environment, and continue from the last point we can reproduce.

For something as dynamic as modern cloud infrastructure and AI workloads, this is an extremely static way to run computation.

The more I’ve worked on this problem, the more I’ve come to think that we are conflating two things that should eventually be separate: the lifetime of the computation and the lifetime of the machine running it.

And as a spoiler, my experiments eventually proved this point right.

THECOMPUTATIONone logicalPython sessionimportload a datasetDataFramesfit a modeldownload filestrainTHEMACHINECPUprepare dataGPUtrainingCPUanalysislarger GPUmodel outgrew the firsttime
The computation keeps going while the machine under it changes.

State has gravity

The field of software infrastructure has spent decades making machines less important, almost a commodity of sorts.

Our source code moved into version control systems like git. Environments became reproducible, thanks to Conda and venv. Virtual machines abstracted physical servers, so that anyone can host their app in a couple of clicks. Containerization mechanisms such as Docker made applications much easier to package and move around. Large-scale cluster schedulers, such as Slurm or PBS, turned fleets of servers into resource pools.

Cloud computing is probably the most extreme example of this concept in the wild - machines are things we provision and destroy whenever we need/want them.

Even live migration as a concept is not new. In 2005, Clark et al. showed at NSDI that a running virtual machine could be moved between physical hosts while keeping downtime very small [1]. The idea that an application should not permanently belong to one physical server is now completely ordinary.

What stopped belonging to one machineWhat made it portable
Source codeversion control, git
EnvironmentsConda, venv
Physical serversvirtual machines
Applicationscontainers, Docker
Fleets of serverscluster schedulers, Slurm and PBS
The machine itselfcloud computing
A running virtual machinelive migration, NSDI 2005
A long-running Python processnot yet

Yet a long-running Python process still develops an enormous amount of attachment to the machine on which it started.

That attachment becomes stronger with time. A ten-second Python process is disposable; a three-hour interactive session might contain a lot of work that exists nowhere else in quite the same form. In more complex scenarios, a "normal" AI workload can take days to complete nowadays, especially given weeks and months to train huge LLMs,

The code of the workload may be reproducible, but reproducing the state can still mean loading data again, rerunning preprocessing, rebuilding indexes, recovering model state, reinstalling dependencies, or simply remembering which sequence of operations produced the thing currently sitting in memory.

You can think of this as state developing gravity. The longer a process runs, the more expensive it becomes to move away from the machine underneath.

cost of moving away from the machinehow long the process has been runningILLUSTRATIVEa ten-second processdisposablea three-hour interactive sessionwork that exists nowhere elsea days-long AI workload

This comes from an observation that, in most data science workloads, the machine that was right when the process started is often not the machine it needs later.

A workload does not need the same machine forever

Let's consider a normal machine learning workflow. Data loading and preprocessing steps might be almost entirely CPU-bound. Training benefits enormously from a good GPU. Evaluation may go back to being mostly CPU work once again. Later, you might run another short burst of training or inference on the GPU.

The infrastructure decision we often make is to allocate around the most demanding phase. If the workflow needs an A100 at some point, it is tempting to put the whole thing on an A100 and leave it there.

idle GPUidle GPUdata loading& preprocessingtrainingevaluationshort burston the GPUGPUCPUallocated: an A100 for the whole workflowneeded
What each phase needs, against an A100 held for the whole run.

At cluster scale, researchers have been attacking this inefficiency for years. Microsoft’s Gandiva showed that dynamically time-slicing and migrating deep-learning jobs could substantially improve utilization in GPU clusters [2]. Pollux, which won a Best Paper award at OSDI 2021, continually adapts how many resources a training job receives based on useful training progress, moving away from a fixed allocation model [3]. Microsoft’s Singularity had a similar idea at a huge production scale, making preemption, migration, and elastic resizing first-class parts of running AI workloads [4].

The underlying lesson is fairly consistent across many of these systems: the correct resource allocation is not necessarily known when a workload begins, and it does not necessarily stay correct.

Real production data tells the same story. A large-scale study of more than 6,000 GPUs at Alibaba found low utilization, load imbalance across heterogeneous hardware, and workloads with very different resource requirements [5]. The problem is not simply that GPUs are expensive. Real workloads are irregular.

Interactive workloads make the mismatch even more obvious. In production notebook traces studied by researchers from Adobe, reserved GPUs were idle for more than 81% of their lifetime, and roughly 3/4 of sessions used their GPU for at most 5% of the session [9]. Based on our production measurements of Clusy, we find very similar results.

ONE RESERVED GPU, ITS WHOLE LIFETIME>81%of it idlebusyidle100 SESSIONS~3/4used their GPUfor at most 5%of the session
Reserved GPUs in production notebook traces. [9]

That is a fairly extreme mismatch between the lifetime of a session and the lifetime of the expensive resource assigned to it.

The obvious answer is to use a GPU only when the computation actually needs one. The less obvious question is: what happens to everything else when you take the GPU away?

What if the session were the durable object?

This is the abstraction we started experimenting with at Clusy.

Instead of taking the machine as the durable object and the Python process as something living inside it, treat the session itself as durable. The CPU, GPU, container, VM, or cloud underneath it becomes fully replaceable.

A Python session could prepare data on a CPU, move to a GPU for training, return to a CPU for analysis, and later move to a larger GPU if the model outgrows the original one, across providers, across architectures, across regions. The user keeps working with the same logical session while the infrastructure underneath it changes, even drastically.

This gives you a slightly different way to think about scheduling. Hardware becomes a phase decision, rather than something you decide once at the beginning of a job.

It also changes how you interpret failures.

Take one of the most familiar messages in machine learning:

CUDA out of memory

Normally that means some combination of lowering the batch size, changing the workload, restoring a checkpoint, or restarting on another machine.

But an OOM contains a simpler piece of information: this workload no longer fits on this device. That does not necessarily imply that the computation itself should disappear.

In experiments with the migration layer we have been building, we let a GPT-2 training session run on a T4 until its memory requirements exceeded the device, then moved the live session to an A100 and continued from the existing state. We tested the move after very different amounts of completed training. We ran a similar experiment with a vision workload that progressed through T4, L4, and A100 hardware as its requirements increased.

GPU memory the workload needstraining progressT4L4A100CUDA out of memory on the T4session moves to an L4outgrows the L4session moves to an A100
The vision run: each out-of-memory moves the live session to a bigger GPU, and training continues from the state it had.

Once you see an OOM this way, it starts looking less like a fatal error and more like a scheduler signal. That distinction is subtle, but I think it points toward a much better programming model.

Fluidity of underneath compute gives tremendous capabilities in terms of scheduling and cost reduction, and provider independence.

Checkpointing solved part of this problem

Of course, the industry already has an answer for preserving expensive computation: checkpoints.

They are indispensable. At Meta, Check-N-Run was designed to checkpoint enormous recommendation models where the ordinary approach was constrained by network bandwidth. On production-scale models, it reduced the write bandwidth needed for checkpointing by 6–17× [6].

There has also been work on avoiding the cost of constant checkpointing altogether. Microsoft researchers explored just-in-time checkpointing, where the system creates a checkpoint in response to evidence of an imminent failure [8].

At a lower level, recent systems work has pushed GPU checkpointing itself forward. PhoenixOS, published at SOSP 2025, explores concurrent checkpoint and restore for GPU applications [10].

All of these systems attack an important problem we care about: losing expensive computational progress is unacceptable.

But there is still a conceptual difference between saving enough application state to recover a job and making a live computational session independent of the machine underneath it.

A Python session is not simply a model checkpoint. It can contain arbitrary Python objects, NumPy arrays, tensor views, files, preprocessors, random-number generators, package state, and references between objects. Some of those relationships matter just as much as the values themselves.

An optimizer pointing at a copy of a tensor instead of the parameter the model is actually using can look superficially correct while being semantically broken. Two tensors might contain the right values but no longer share the same underlying storage. A restored random-number generator might produce a different training trajectory.

RELATIONSHIPS SURVIVESUPERFICIALLY CORRECTsemantically brokenoptimizer and modelmodel.weightoptimizera tensor and its viewab = a[::2]storageoptimizer and modelmodel.weightoptimizercopy of weightsame valuesa tensor and its viewab = a[::2]storagestorage, again
The same values either way. Only one of these keeps training.

Even hardware introduces ambiguity. Move exactly the same stored tensor values from one GPU architecture to another and subsequent floating-point calculations can legitimately differ slightly.

Preserving state and reproducing arithmetic are related problems, but they are not the same problem. This is why moving a live Python process becomes much more complicated and quite more interesting than calling pickle.dump().

What we built at Clusy

Over the last few months, we have been building this abstraction underneath Clusy, our environment for data science and computational work.

We call it SOMA - Session-Oriented Migration Abstraction.

(Also, it's the name of the district and motel in SF where we lived for a while)

The high-level idea is simple: a logical Python session has an identity that is independent of the runtime currently executing it. The low-level implementation is months of work and tonnes of filtered research.

We capture the supported application state, files, and environment information into a portable representation called a capsule, reconstruct it at the destination, validate the result, and only then transfer execution authority.

The migrations are done atomically, so until a move commits, the source remains the authoritative copy of the session. If restoration fails halfway through, the destination can be discarded, and the original session continues serving and can be reattempted.

capturecapsulereconstructvalidatetransfer execution authoritySOURCEapplication statefilesenvironmentauthoritative until commitDESTINATIONapplication statefilesenvironmenttakes over at commita portable capsulerestore fails halfway: discard the destination, the source keeps serving
Nothing changes hands until the destination has been validated.

This migration can happen across completely different hardware, isolation abstractions (containers, VMs, processes), providers and regions.

We deployed this in notebooks because notebooks make the problem very visible. They are long-lived, interactive, and stateful by design: you load things into memory, inspect outputs, train a model, modify something, execute again, and gradually turn the kernel into a working environment.

But the abstraction itself is not particularly notebook-specific. The broader idea is simply that a Python computation should be able to outlive a runtime.

Today on Clusy, a CPU session can run in one style of sandbox and later execute on GPU infrastructure with a different provider, isolation layer, base image, and device. The underlying state mechanism has already been used across more than 10,000 production sessions on the platform.

In our experiments, moving between appropriate CPU and GPU phases reduced the cost of some workflows by as much as 45–70%. We were also able to keep training state intact while changing chips without reconstruction.

10,000+

production sessions have used the underlying state mechanism

45–70%

lower cost on some workflows, moving between CPU and GPU phases

SOMA in production on Clusy, and in our experiments.

None of this should ultimately feel impressive to the person using the product. The ideal interface is actually quite boring: you have a Python session, and the compute underneath it changes when necessary.

That is the part we care about.

The cloud is already becoming fungible

There is a broader infrastructure trend here.

Some of my friends at Berkeley developed SkyPilot, which asks users to think less about individual clouds and more about a pool of resources spanning providers [7]. Instead of deciding that a workload “belongs” to AWS, GCP, Azure, or some other provider, the system can choose between them based on cost, availability, and hardware requirements.

The common direction is that infrastructure is becoming increasingly fungible.

If the state can move too, the scheduler gets a much larger design space.

A CPU can be temporary. A GPU can be temporary. A cloud provider can be temporary. Eventually, perhaps even the distinction between local and remote execution becomes much less important than it is today. In fact, we have already extended our mechanism to bridge this gap and will soon deploy the updated version.

The logical computation is what remains.

Agents make this more interesting

I think this becomes especially important as agents start controlling longer-running computational workflows.

Humans have accumulated a lot of infrastructure rituals. We choose machine types, watch memory usage, save checkpoints, release GPUs, restart processes, and decide when something needs to move to a bigger box. Those are useful skills under today’s abstraction, but there is no reason an agent needs to inherit all of them.

An agent can observe that the current phase is CPU-bound and release the GPU. It can notice increasing memory pressure and choose a larger accelerator before the job fails. It can respond to a preemption notice, trade cost against latency, or move away from unavailable capacity.

The choice of runtime becomes one more decision inside the agent’s control loop and tooling.

The important shift is that the agent does not need to think in terms of “migration.” It can simply ask:

What is the best place to execute the next operation?

Your source code is already portable. Increasingly, your environment is portable. The next step is making the live computation portable as well.

The machine is not the computation. It is just where the computation happens to be running right now.

We push towards fluid compute.

Follow how we do it at @clusyio

References

[1] Clark, C. et al. “Live Migration of Virtual Machines.” NSDI 2005.

[2] Xiao, W. et al. “Gandiva: Introspective Cluster Scheduling for Deep Learning.” OSDI 2018.

[3] Qiao, A. et al. “Pollux: Co-adaptive Cluster Scheduling for Goodput-Optimized Deep Learning.” OSDI 2021.

[4] Shukla, D. et al. “Singularity: Planet-Scale, Preemptive and Elastic Scheduling of AI Workloads.” EuroSys 2022. Microsoft.

[5] Weng, Q. et al. “MLaaS in the Wild: Workload Analysis and Scheduling in Large-Scale Heterogeneous GPU Clusters.” NSDI 2022.

[6] Eisenman, A. et al. “Check-N-Run: a Checkpointing System for Training Deep Learning Recommendation Models.” NSDI 2022.

[7] Yang, Z. et al. “SkyPilot: An Intercloud Broker for Sky Computing.” NSDI 2023. UC Berkeley.

[8] Gupta, T. et al. “Just-In-Time Checkpointing: Low Cost Error Recovery from Deep Learning Training Failures.” EuroSys 2024. Microsoft Research.

[9] Carver, B. et al. “NotebookOS: A Replicated Notebook Platform for Interactive Training with On-Demand GPUs.” ASPLOS 2026. Adobe Research, George Mason University, University of Virginia.

[10] Wei, X. et al. “PhoenixOS: Concurrent OS-level GPU Checkpoint and Restore with Validated Speculation.” SOSP 2025.