Data Science as a Search Problem
The next experiment depends on the last result. Branchable notebook state gives data science agents more than one path to explore.
Most data science questions don’t really have a next step. They have several.
Working with these workflows throughout my academic and professional career, I came to this realization fairly early.
Sure, you clean a dataset, look at a few distributions, build a baseline, and then things get interesting. Maybe XGBoost is the obvious next model. But so is CatBoost. Maybe the model is fine, and the representation is wrong. Maybe you should remove a suspicious feature, change the split, optimize a different objective, or simply check whether the result survives another random seed.
In practice, we usually pick one direction and keep going.
Run something, wait, inspect the result, modify the code, run again. If we want to seriously explore several directions at once, it gets more annoying. We can write parallel code, submit several jobs, create a workflow, or just copy the notebook a few times and try to remember what changed where.
None of this is technically impossible. Parallel computers have existed for a very long time.
What feels outdated is that the investigation itself is still usually represented as one linear object.
Look at it another way: almost every sufficiently open-ended data science problem is a search problem.
The experiment is already a search tree
Take the current state of an analysis and imagine it as a node.
Everything you know so far lives there: the cleaned dataset, the features you created, the model you trained, the plots you looked at, the assumptions you currently believe.
From that node, there are several possible actions. Try a different model. Transform the data. Remove a feature. Change the loss. Collect more evidence.
Each action creates another state.
Now you have branches.
Some branches fail immediately. Others look promising and deserve more compute. One surprising result might cause you to return to an earlier state and explore an entirely different direction.
If you keep going, what you are doing starts to look a lot like a search tree.
You don’t know the tree beforehand. It emerges while you work.
This differs from a fixed workflow DAG where the computation is specified before execution. In research, the outputs determine which computation should exist next.
This isn't a particularly radical description of science. We form hypotheses, test them, discard some and develop others. What is interesting is how much of machine learning already formalizes parts of this idea.
The original Google Vizier paper opens with an observation I really like: sufficiently complex systems can become easier to experiment with than to understand [1]. That describes a surprising amount of modern ML. We usually don’t know the correct architecture, features, preprocessing strategy, or hyperparameters before we start. If we did, there would not be much searching to do.
Bergstra and Bengio showed years ago that random search could outperform grid search for hyperparameter optimization because only a subset of hyperparameters tend to matter strongly for a given problem [2]. Similarly, Hyperband suggested starting with many configurations, spending a little compute on each, killing weak candidates early, and allocating more resources to the promising ones. It reported more than an order-of-magnitude speedup over competing approaches on some benchmarks [3].
In one of my favorite papers, Microsoft described a similar pattern from a systems perspective in Gandiva. They called deep learning “feedback-driven exploration”: researchers launch related experiments, inspect early feedback, and then change where their resources go. In their experiments, Gandiva accelerated hyperparameter searches by up to an order of magnitude [4].
So the search itself is not a new idea.
What I find weird is that most of our everyday interactive tools still make the search look like a line.
We flattened the search tree into a document
Notebooks make this especially visible.
I like notebooks. We built Clusy around them for a reason. They are still one of the dominant interfaces for interactive ML work: in the 2024 Python Developers Survey, 50% of respondents who trained or generated predictions with ML models reported using Jupyter Notebook as a training platform [5]. Respondents could select more than one platform.
And the basic interaction is excellent for research. Write code, execute it, inspect the result, think for a bit, change something, and continue.
The awkward part appears when one state has several reasonable futures.
Imagine you’ve spent half an hour loading data, cleaning it, constructing features, and getting the environment exactly where you want it. Now there are three things worth testing.
One approach changes the representation and trains CatBoost. Another keeps the current features and tries XGBoost. A third tries a small neural network.
All three begin from exactly the same state.
In a normal notebook, you can put them further down the page and execute them sequentially. You can turn them into separate jobs and reconstruct the shared setup. Or, very realistically, you duplicate the notebook.
Anyone who has done enough research has seen some version of:
experiment.ipynb
experiment_new.ipynb
experiment_new2.ipynb
experiment_final.ipynb
This isn’t really a criticism of Jupyter. Chattopadhyay and colleagues documented how exploratory notebook work makes it difficult to follow execution history, compare notebook versions, and recover earlier analyses [6].
The deeper issue is that the interface gives you one mutable present.
But research frequently needs several possible futures.
What if experiments could actually fork?
Suppose the state of your analysis were branchable.
You reach that useful point after preprocessing and simply fork it three ways. Each branch starts from the same prepared state, then diverges.
They run at the same time.
Twenty minutes later, the neural network is obviously worse, so you prune that branch. XGBoost looks surprisingly good, so you expand it: one branch changes the feature set, another changes the objective, another tests several seeds.
Maybe one of those results reveals leakage. Now the interesting move is to revisit the preparation, fix the split, and create a new branch from the corrected state. In an ideal environment, returning to an earlier state would be a first-class operation; otherwise, you have to reconstruct it.
That is a search process.
The key point isn't that everything must become a visible graph on screen. I actually think that would become annoying very quickly.
The graph is a better execution model.
The notebook can remain the interface you work in. Python can remain Python. The difference is that a state no longer has to have exactly one successor.
One point in an investigation can have many futures. And once that is possible, parallel computing becomes useful at a much higher level.
We are no longer just parallelizing matrix operations or making one training run faster. We are parallelizing the search.
Why did we make the search sequential in the first place?
Mostly because managing parallel experiments is annoying.
If I have five plausible hypotheses, I can theoretically launch all five today. But then I need to keep track of five environments, five pieces of code, five sets of outputs, which state they started from, which ones failed, and what should happen next.
At some point, sequential execution becomes easier simply because I am the scheduler.
This is one reason agents make the idea much more interesting.
An agent could inspect four results as they arrive, compare them against the objective, stop spending compute on two branches, expand another, and revisit the preparation if something looks wrong.
The shape of the problem starts to resemble an actual search algorithm: maintain candidate states, evaluate them, expand the useful ones, prune the bad ones.
You don't need to literally run A* over your notebook. The analogy only needs to go so far.
But I think the mental model is important.
Instead of asking an agent:
What should I try next?
we can eventually ask:
What parts of this search space are worth exploring?
There is an interesting parallel in LLM research. Tree of Thoughts explored what happens when a model can maintain and evaluate several reasoning paths instead of committing immediately to one chain. On 100 relatively hard Game of 24 puzzles, the paper reported an average 4% success rate for GPT-4 with chain-of-thought prompting and 74% for Tree of Thoughts using breadth-first search with five candidate states [7]. The search used more inference work; this was not a comparison at equal compute.
That is not to say that reasoning benchmarks and data science experiments are the same thing.
But the underlying intuition is useful: if you don't know which path is correct, committing to one path early throws information away.
In data science, those branches can be much more concrete. They are not hypothetical thoughts. They can be real Python processes running real experiments against real data.
This is something we’re building into Clusy
This is one of the ideas we've been working on directly in Clusy.
From a notebook branch’s current state, you can create child branches with their own kernels and working files. They inherit supported in-memory Python state, run different code concurrently, and can themselves become starting points for further branches.
Today, those branches share the project’s compute resources. State transfer has limits: live connections and other unsupported objects may need to be recreated. Forking uses the current state; it does not restore an arbitrary earlier cell.
The broader direction is to help agents compare the evidence and decide which branches deserve more work. Automatic, metric-guided expansion and pruning are still part of that research direction.
We aren't trying to turn the notebook into a giant visual workflow editor. I still want to open a notebook and write Python.
The difference is underneath.
Today, an agent working in a normal environment will often behave sequentially: propose one approach, execute it, inspect what happened, modify it, and continue.
With branchable execution, several approaches can be reasonable enough to test before committing to one. There is still a compute budget, and a good search policy has to decide where another experiment is worth its cost.
Then let the results determine where the search goes next.
To us, that starts feeling much closer to how a useful research agent should behave.
The question is the search space
The more I think about this, the less convinced I am that a notebook, script, or training job is really the natural unit of data science.
The thing we actually care about is the question.
Around that question is a search space. We already create this structure manually. It lives across notebook copies, experiment trackers, Git commits, cluster jobs and, mostly, our own memory. Maybe that structure should simply be part of the computational model?
Not every analysis needs a search tree. Sometimes there genuinely is one obvious next step, and introducing branching would only make the work more complicated.
But the interesting problems are usually the ones where we don't know the answer yet.
And for those, I think there is a useful shift in perspective:
A data science problem is not just a program to execute.
It is a space to search.
References
[1] Golovin, D. et al. Google Vizier: A Service for Black-Box Optimization. KDD, 2017.
[2] Bergstra, J. and Bengio, Y. Random Search for Hyper-Parameter Optimization. JMLR, 2012.
[3] Li, L. et al. Hyperband: A Novel Bandit-Based Approach to Hyperparameter Optimization. JMLR, 2018.
[4] Xiao, W. et al. Gandiva: Introspective Cluster Scheduling for Deep Learning. OSDI, 2018.
[5] Python Software Foundation and JetBrains. Python Developers Survey 2024. Training-platform question, respondents who train or generate predictions with ML models.
[6] Chattopadhyay, S. et al. What’s Wrong with Computational Notebooks? Pain Points, Needs, and Design Opportunities. CHI, 2020.
[7] Yao, S. et al. Tree of Thoughts: Deliberate Problem Solving with Large Language Models. NeurIPS, 2023.