Blog

Data Science Needs More Than a Coding Agent

The deeper problem is that data science isn't primarily a code generation problem.

Research10 min readEldar Hasanov

Coding agents are having a pretty incredible couple of years.

Claude Code, Codex, Cursor, and others have become default tools in so many development workflows. I personally do not remember a time when I did not use any of them for at least parts of a project in the past 2-3 years.

In JetBrains’ 2026 developer survey, 90% of professional developers said they use coding agents at least weekly, and 68% use them daily.

And unsurprisingly so, because they have gotten genuinely good.

On JetBrains’ recent Kotlin benchmark, the best agent setup resolved 90 of 105 real coding tasks (85.7%). On famous SWE-bench Verified, frontier agents are already reaching 80.9% of real GitHub issues resolved, while newer SWE-bench variants are already pushing toward saturation.

In fact, models tackle these benchmarks so fast that there is an entire research field of building new benchmarks to test intelligence more deeply.

But there’s one area where the experience still feels surprisingly different:

Data Science.

Give an agent a well-scoped software task (e.g., "fix this bug", "implement this function", "refactor this module") and the workflow is increasingly reliable.

Now give it a dataset and say:

Figure out what matters here. Clean it, explore it, build a model, evaluate it, and tell me what you find.

Surprisingly, things get much harder. Not because frontier models can't write pandas or PyTorch. They obviously can.

The deeper problem is that data science isn't primarily a code generation problem.

A real data science workflow is a loop

A normal task rarely looks like: requirement → code → done

It looks more like:

look at the data → form an idea → form hypothesis → write code → execute it → inspect what happened → observe live patterns → change your mind → execute again

A NORMAL TASKrarely looks likerequirementcodedoneA REAL DATA SCIENCE WORKFLOWlook at the dataform an ideaform hypothesiswrite codeexecute itinspect what happenedobserve live patternschange your mindexecute againthe output of one computationbecomes the inputto the next decision
  • Maybe a column turns out to be 40% missing.
  • Maybe the classes are massively imbalanced.
  • Maybe validation score looks suspiciously good, and you discover leakage.
  • Maybe XGBoost gives you 0.84 AUC, so you try CatBoost.
  • Maybe your neural network is overfitting, so you change the augmentation strategy.
  • Maybe the plot you just produced makes the entire hypothesis look wrong.

The output of one computation becomes the input to the next decision.

That's fundamentally different from just producing correct code. The problem rarely has a deterministic outcome that you can write a test for, and the best possible solution is unknown by default. It is a space of exploration within a pipeline-like structure

And we're finally starting to get benchmarks that actually measure it.

Frontier agents still struggle with end-to-end data science

This statement originated from my personal experience during my research projects at NASA, USC, Berkeley, and Imperial College.

As a software engineer and data-focused researcher, I have always observed this gap between using mighty coding agents to help me build systems, platforms, and applications, and seeing those same models struggle with even simpler data tasks in my research.

A paper released last month, DSAgentBench, is probably the clearest example I've seen out there. Instead of asking models isolated Python questions, the researchers built 275 end-to-end data science tasks spanning:

  • data wrangling
  • exploratory analysis
  • modeling
  • visualization
  • validation

The agents have to operate in real computer environments using things like notebooks, terminals, IDEs, browsers, and databases. More importantly, they have to make decisions based on the outputs they generate along the way.

Across 15 models, the strongest agent tested completed only 56.7% of the tasks successfully.

The researchers specifically point to failures in tool orchestration, operating-system grounding, and multi-step reasoning.

Although it's unfair to directly compare this to 80-90%+ on SWE benchmarks, as those are different.

But the difference in what they're testing is interesting.

One asks whether an agent can solve software-engineering tasks.

The other asks whether an agent can conduct a computational investigation.

That second problem is much messier.

Best reported result, share of tasks

Solve software-engineering tasks

JetBrains Kotlin benchmark

best agent setup · 90 of 105 tasks

85.7%

SWE-bench Verified

frontier agents · real GitHub issues

80.9%

Conduct a computational investigation

DSAgentBench

strongest of 15 models · 275 end-to-end tasks

56.7%

Different benchmarks testing different things, so this is not a like-for-like comparison.

Sources [2] [3]

ML experimentation makes this even more obvious

OpenAI's MLE-bench takes the idea further. ML nowadays is one of the most common and "hype" data science workloads we could look at.

Instead of synthetic programming problems, MLE-bench gives agents 75 real Kaggle competitions. The agent gets data and compute and has to do the work: prepare datasets, train models, and run experiments.

When the benchmark was introduced in 2024, the strongest configuration tested (o1-preview with the AIDE agent scaffold) reached at least Kaggle bronze-medal performance on only 16.9% of competitions.

Models have obviously improved dramatically since 2024, so I wouldn't use 16.9% as a statement about today's frontier.

Today, the best comparable result using heavy-hitter models with all the right configurations on the official leaderboard is 64.4%. But the difficulty curve is quite revealing. On the benchmark's hardest group of competitions, even the strongest recent systems reach medal-level performance on only around 42–47% of tasks.

MLE-bench · medal level, of 75 Kaggle competitions

2024, at launch

o1-preview with the AIDE scaffold · bronze or better

16.9%

Today, best on the leaderboard

heavy-hitter models, tuned configurations

64.4%

Today, hardest competitions

strongest recent systems

42–47%
Source [4]

The important part is the kind of problem being measured. To succeed, the agent can't simply know how to write a training loop. It has to decide which training loop is worth trying.

Then run it.
Then interpret the result.
Then decide what to try next.

That's much closer to what data scientists actually spend their time doing.

And this matters because notebooks are everywhere in data science

For a huge part of the data world, this experimentation happens inside a notebook. Jupyter and Jupyter-compatible notebooks are still one of the core interfaces for exploratory analysis, modeling, and research.

One JetBrains survey found Jupyter usage as high as 70% among respondents doing activities like exploratory data analysis, data visualization and ML modeling. Another found roughly half of data science professionals using Jupyter notebooks to present their work.

Google said Colab, itself built around the Jupyter notebook model, had already passed 10 million users by 2023, with reported 1-3 million users per day today.

And notebook usage is increasingly hidden inside other tools.

People might say their main editor is VS Code or PyCharm while still spending significant time working with .ipynb files and Jupyter kernels inside those environments. In the 2024 Python Developer Survey, for example, 53% of VS Code users reported using its Jupyter support.

70%

Jupyter usage at its highest, among people doing EDA, visualization and ML modeling

~½

of data science professionals present their work in Jupyter notebooks

10M+

Colab users by 2023, with a reported 1–3M a day today

53%

of VS Code users use its Jupyter support

JetBrains surveys and Google's Colab numbers. [5] [6]

Notebooks work so well for this kind of work for a reason.

They put:

code + execution + outputs + plots + explanation

in the same place.

But that creates a strange problem for agents.

A notebook isn't just a collection of files

A traditional coding agent can understand a lot about a software project by looking at the filesystem.

There is a repository.
There are files.
There are tests.

Change the files, run the tests, inspect what failed, repeat.

In data science, there is another layer:

runtime state.

  • Your dataframe may already be sitting in memory.
  • A model may have been trained three cells ago.
  • Cell 8 may depend on Cell 3 having been executed, even if the visible notebook doesn't make that obvious.
  • There may be a GPU process running.
  • There may be a 20GB dataset mounted somewhere.

The most important piece of context might not be in the source code at all.

It might be the plot the previous cell just produced.

THE FILESYSTEMrepo/analysis.ipynbfeatures.pytests/another layerRUNTIME STATEdfalready in memorymodeltrained three cells agoCell 8 → Cell 3an invisible dependencyGPU processstill running20GB datasetmounted somewherethe last plotjust producedcoding agentreads the filesdata science agentneeds more than files
Files on top. Runtime state underneath.

That means an effective data science agent needs more than access to files.

It needs to understand and operate the computational environment itself.

The model is only part of the system

Hugging Face published a great example of this while building its own Jupyter Agent.

Their goal was to let an LLM actually execute code inside notebooks and use those executions to solve data-analysis problems.

At one point, they took the same small Qwen model and changed the agent scaffolding around it.

On the easy portion of their benchmark, accuracy went from:

44.4% → 59.7%

The model didn't suddenly get smarter. It's the environment around the model that got better.

I think this is an underappreciated part of the current agent wave. We spend enormous amounts of time comparing:

Claude vs GPT vs Gemini

But as agents become more capable, another question becomes increasingly important:

What environment are we putting the model inside?

Give a great model poor context, brittle tools, and no useful execution feedback, and it will still fail.

Give it an environment designed around the actual task, and the exact same model can behave very differently.

I think data science needs its own agent infrastructure

Software-engineering agents increasingly have a very good environment.

They can search repositories, edit files, run commands, execute tests, inspect errors, use Git, open pull requests, and work in parallel.

The equivalent environment for data science is still being figured out.

A serious data-science agent probably needs to be able to:

understand the live runtime, manage compute, execute code, inspect outputs, work with large datasets, control compute, recover from failed experiments, preserve useful state, and compare multiple approaches.

And eventually, I think one of the biggest changes will be parallelism.

Due to the linear nature of notebooks, people usually try experiments sequentially:

Try XGBoost.
Wait and see.
Change the parameters.
Wait and see.
Try LightGBM.
Wait and see.

That's partly because a human can only actively drive so many experiments at once.

Agents don't have the same constraint.

If there are four reasonable hypotheses, why shouldn't an agent branch the environment, try all four, inspect the results, and continue from the promising ones? And I mean proper branch, not manage dozens of notebooks doing parallel things, repeating the same preprocessing steps. Handling N amount of state for N notebooks is not a real solution for either agents or humans.

SEQUENTIALone at a timeXGBoostwait and seechange the parameterswait and seeLightGBMwait and seeBRANCHEDfour hypothesespreprocessingXGBoostLightGBMCatBoostneural netcontinue fromthe promising onestime
One at a time, or branched from the same preprocessed state.

The workflow starts looking less like autocomplete and more like a search process over experiments.

This is also why we're building Clusy

When we started working on @clusyio, the obvious version of an “AI notebook” was basically:

Jupyter + chat.

The more we built, the less interesting that seemed.

If an agent is going to do meaningful data science, it shouldn't just sit beside the notebook and suggest code. It needs to operate the environment.

It should be able to run the analysis, see the result, manage the kernel, install something when it needs it, move to different compute, search a research paper with modern approaches, branch an experiment, preserve state, and keep going.

We're still early, and there are plenty of hard problems here.

But I increasingly think the gap between today's coding agents and a genuinely useful AI data scientist isn't going to be closed by model intelligence alone, at least not fast enough.

The models will keep getting better.

The bigger question is whether the environments around them catch up. Whatever tooling, harnesses, and runtimes are built around these models will only improve in an amplified manner as they get smarter.

Till maybe one day, there is a specialized model for data science? But on that, some time later.

Because data science isn't just about writing the right code. It's about running something, looking at what happened, and deciding what to do next.

And that's exactly the kind of loop agents should eventually be very good at.

We are on a mission to explore this question further. Join us on our journey at https://www.clusy.io/

Sources

[1] JetBrains — AI Coding Agents: Adoption Trends, Developer Ecosystem Survey 2026.

[2] JetBrains — Kotlin Benchmark for AI Coding Agents, 2026.

[3] Rahman et al. — DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments?, 2026.

[4] OpenAI — MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering, 2024.

[5] JetBrains — Developer Ecosystem surveys on Jupyter usage in data science.

[6] Google Research — Colab usage, 2023 Year in Review.

[7] Hugging Face — Jupyter Agents: training LLMs to reason with notebooks, 2025.