Blog

The Benchmark Gap: Coding Agents vs Data Science

Coding agents got very good partly because we built a strong execution and verification loop around the model. For data science, we are still figuring out what that loop should be.

Research5 min readEldar Hasanov

We are a generation of engineers who were lucky (or unlucky) to witness this unprecedented speed of improvements in models for software engineering.

Although software can be considered one of the most "cutting-edge" industries that changes fast and often, one could argue that this has probably been the most monumental shift in the field so far, and it is still going.

Of course, this improvement is not just a feeling; it is backed by recognized benchmark numbers for various coding agents and clearly tracked.

1. Software engineering benchmarks are getting very high

One of my favorite IDE companies, JetBrains, recently built a benchmark from 105 real Kotlin engineering tasks taken from active open-source repositories. Claude Code with Opus 4.7 solved 90 of them (85.7%). Junie and Codex both landed around 82% [1]. Keep in mind, this was a couple of Opuses ago; at this point, we are waiting for updated figures.

SWE-bench Verified had a similar trajectory. Frontier performance climbed from 74.9% to 80.9% in six months before OpenAI stopped using it for frontier evaluation. The reason was almost the opposite of what you might expect: they found serious test flaws and increasing contamination, meaning the benchmark was no longer clean enough to tell whether further gains represented better agents or prior exposure to the tasks [2].

That led to harder replacements. SWE-Bench Pro, published at ICML 2026, contains 1,865 problems from 41 active repositories and is explicitly designed around longer, more realistic engineering work [3]. The fact that the benchmark ecosystem itself has had to keep moving toward harder tasks is a useful signal about how quickly coding agents have progressed.

It is not surprising that some of the early classic benchmarks we used are now basically saturated; the tasks now seem ridiculously easy with current capabilities.

It's not surprising that people are actually using these tools. JetBrains' 2026 survey of more than 15,000 professional developers found 90% using coding agents at least weekly and 68% daily [4].

Software engineering is obviously not exactly solved, but some would argue it's almost solved. But handing a reasonably scoped repository task to an agent is no longer an experiment. It is normal work.

Interestingly, data science tasks tell a completely different story.

2. The data science benchmarks are still much rougher

DSAgentBench is probably the cleanest example. It contains 275 end-to-end tasks across data wrangling, exploration, modeling, visualization, and validation, with agents operating actual notebooks, terminals, IDEs, and databases [5].

The strongest tested agent completed only 56.7% of tasks. The authors specifically identify tool orchestration, OS grounding, and multi-step reasoning as persistent failure modes.

As another example, OpenAI's MLE-bench pushes closer to real ML engineering. Agents get 75 Kaggle competitions and have to prepare the data, train models, and improve their submissions. [6]

The progress here has been pretty good. The original best result in 2024 was about 17%. The latest comparable result on OpenAI's official leaderboard reaches 64.4% medal-level performance overall. But on the high-complexity competitions, that same system reaches only 42.2%. OpenAI paused new leaderboard submissions in April 2026 while it works on keeping comparisons fair.

Let's move from ML engineering into more complex AI/ML research, and watch the gap widen even more.

RExBench, published at ACL 2026, asks coding agents to implement realistic extensions to existing AI research papers. Every tested agent failed the majority of tasks. The best system reached roughly 33%, and even with human-written hints the best performance remained below 44% [7].

OpenAI's PaperBench is harsher still. Agents are given 20 ICML 2024 Spotlight and Oral papers and asked to replicate them from scratch: understand the contribution, write the implementation, and reproduce the experiments. The best tested agent achieved an average replication score of 21.0% and did not beat the ML PhD baseline [8].

BEST REPORTED RESULTnot directly comparable: different task distributions0%25%50%75%100%85.7%KotlinJetBrains80.9%SWE-benchVerified64.4%MLE-benchoverall56.7%DSAgentBenchend to end42.2%MLE-benchhigh complexity~33%RExBenchextensions21.0%PaperBenchreplicationwith hints,still under 44%software engineeringdata science · ML engineeringAI/ML research
Best reported results from the benchmarks in this post, ordered by the kind of work. [1] [2] [5] [6] [7] [8]

These benchmarks are not directly comparable. An 85.7% Kotlin score and a 56.7% DSAgentBench score do not imply a 29-point capability gap; they measure completely different task distributions.

What I find interesting is the pattern across them instead, and how the speed of improvements and even the standing of current capabilities is so different.

Frontier models clearly know how to write Python. That is probably not the bottleneck.

Software engineering gives the agent a very clean feedback loop: inspect the repository, modify something, run the tests, see what failed, and try again. There is usually some external definition of correctness that agents and humans all agree on in most cases.

Data science often doesn't have that.

The model might execute perfectly valid code and still reach the wrong conclusion. A suspicious validation score, leakage in a feature, the wrong train/test split, an odd distribution, or a training curve can matter more than whether the program successfully ran.

Running the code is often where the reasoning starts.

SOFTWARE ENGINEERINGDATA SCIENCEinspect the repositorymodify somethingrun the testssee what faileddecide what to trustwhat to try nextrun the coderead the outputstests?an external definition of correctnessvalidation scoreleakagedistributiontrain/test splittraining curve
The same loop twice: one with an agreed judge in the middle, one where the evidence has to decide.

That changes what the agent needs around it. It has to understand intermediate state and outputs, compare experiments, manage compute, decide what evidence to trust, and determine what is worth trying next.

This is a lot of what we think about while building Clusy. We started with notebooks because code, live state, and experimental outputs naturally meet there, but I think the underlying problem is much broader than notebooks.

Coding agents got very good partly because we built a strong execution and verification loop around the model.

For data science, we are still figuring out what that loop should be.

And right now, the benchmarks make that pretty obvious.

Follow our journey with @clusyio to see if we can figure it out.

References

[1] Chernyaeva, A. et al. “Introducing the Kotlin Benchmark for AI Coding Agents.” JetBrains, 2026.

[2] OpenAI. “Why SWE-bench Verified No Longer Measures Frontier Coding Capabilities.” OpenAI, 2026.

[3] Deng, X. et al. “SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?” ICML 2026.

[4] Bogdanov, M. et al. “AI Coding Agents: Adoption Trends.” JetBrains Developer Ecosystem Survey, 2026.

[5] Rahman, M. et al. “DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments?” 2026.

[6] Chan, J. S. et al. “MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering.” 2024.

[7] Edwards, N. et al. “RExBench: Can Coding Agents Autonomously Implement AI Research Extensions?” ACL 2026.

[8] Starace, G. et al. “PaperBench: Evaluating AI’s Ability to Replicate AI Research.” 2025.