A benchmark for autonomous data production Research preprint

AutoDataBench

Can agents write the data that feeds the self-improvement loop?

THE UNIT OF EVALUATION
01 / INPUTOriginal task
+ target model attempts
02 / AUTHORAgent writes
one new executable task
03 / ACCEPTRun and check
before any training
One delivered artifact. One evidence-based verdict.
24original tasks
across three suites
5author agents
against one target model
18.4/100highest score
at 45 minutes per task
The central findingAgents can write tasks that recreate decisions the target model struggled with. They struggle to set task difficulty.
01 / THE QUESTION

What would it take to create agentic training tasks without human task authors?

Agentic training depends on executable tasks: an environment, an instruction, and a verifier that can tell whether the work was done. Running a model on those tasks produces the trajectories used for training. Today, people still set the standard and check what a task generator produces.

AutoDataBench isolates one part of that job. An author agent reads an existing task and a record of the target model attempting it, then writes one different executable task for the same suite.

We evaluate that delivery against the kinds of acceptance checks widely used in training-data production: is it valid and new, at a useful difficulty, and aimed at the intended behavior? These checks happen before any training run. A later training score cannot tell us what this particular task contributed to a batch.

Read the longer explanation
02 / THE METHOD

One task in. One new task out. Then three checks.

A usable delivery must satisfy all three requirements. The score is their product, so a task with a disqualifying defect or the wrong difficulty scores zero.

01

Usable and new

A gate rejects defects such as a verifier that never checks the result, an exposed answer, or the original task with only its surface changed.

GATE · 0 OR 1
02

Difficulty in range

The target model attempts the delivered task six times. It must solve the task at least once and at most four times.

DIFFICULTY · 0 OR 1
03

Right behavior

A judge checks the new attempt transcripts against a hidden rubric of behavioral modes, and must cite evidence for each mode it marks present.

QUALITY · 0 TO 1
EPISODE SCOREgate × difficulty × qualityJudged before training
03 / RESULTS

The bottleneck is difficulty calibration.

At the default 45-minute budget, no author agent scores above 20 out of 100. Only 14.6%–25.0% of deliveries land in the target model's usable pass-rate band. Among deliveries that also clear the gate, quality ranges from 0.700 to 0.967.

Scores average two episodes per original task, then average across tasks. The same target model and hidden rubrics are used for every author agent.
OVERALL SCORE

Five author agents

0–100 scale · higher is better
0255075100
WHAT FAILSMost tasks miss the usable difficulty band.

They tend to be solved on every attempt or on none. A task at either end is rejected.

THE SHORTCUTCopying the original can preserve its difficulty.

That is why the gate rejects a surface swap: the agent must give the solver something new to work out.

MORE TIME45 → 180 minutes raises Kimi K3's score from 18.4 to 54.1.

It produces more usable tasks, while the cost per usable task stays roughly level.

View the full result table
Results from the current paper, default 45-minute budget
Author agentIn bandGate passed
given in band
QualityScore
kimi-k321.3%90.0%0.9670.184
gpt-5.6-sol21.3%100.0%0.8330.177
qwen3.8-max14.6%100.0%0.9520.139
glm-5.321.7%90.0%0.7000.130
deepseek-v4-pro25.0%41.7%0.9310.097
04 / THE BENCHMARK

Three suites. Different kinds of work.

Eight original tasks come from each suite. The author agent writes a new task in that suite's own format; the target model and scoring procedure remain fixed.

The ranking changes across suites. No author agent leads everywhere, so the overall score should not stand in for domain-specific performance.

05 / RESOURCES

Read, inspect, or run the benchmark.

The benchmark tests whether a delivered task meets acceptance criteria. Whether tasks that pass these checks improve a model after training remains an open question.