Usable and new
A gate rejects defects such as a verifier that never checks the result, an exposed answer, or the original task with only its surface changed.
GATE · 0 OR 1Can agents write the data that feeds the self-improvement loop?
Agentic training depends on executable tasks: an environment, an instruction, and a verifier that can tell whether the work was done. Running a model on those tasks produces the trajectories used for training. Today, people still set the standard and check what a task generator produces.
AutoDataBench isolates one part of that job. An author agent reads an existing task and a record of the target model attempting it, then writes one different executable task for the same suite.
We evaluate that delivery against the kinds of acceptance checks widely used in training-data production: is it valid and new, at a useful difficulty, and aimed at the intended behavior? These checks happen before any training run. A later training score cannot tell us what this particular task contributed to a batch.
A usable delivery must satisfy all three requirements. The score is their product, so a task with a disqualifying defect or the wrong difficulty scores zero.
A gate rejects defects such as a verifier that never checks the result, an exposed answer, or the original task with only its surface changed.
GATE · 0 OR 1The target model attempts the delivered task six times. It must solve the task at least once and at most four times.
DIFFICULTY · 0 OR 1A judge checks the new attempt transcripts against a hidden rubric of behavioral modes, and must cite evidence for each mode it marks present.
QUALITY · 0 TO 1At the default 45-minute budget, no author agent scores above 20 out of 100. Only 14.6%–25.0% of deliveries land in the target model's usable pass-rate band. Among deliveries that also clear the gate, quality ranges from 0.700 to 0.967.
Scores average two episodes per original task, then average across tasks. The same target model and hidden rubrics are used for every author agent.They tend to be solved on every attempt or on none. A task at either end is rejected.
That is why the gate rejects a surface swap: the agent must give the solver something new to work out.
It produces more usable tasks, while the cost per usable task stays roughly level.
| Author agent | In band | Gate passed given in band | Quality | Score |
|---|---|---|---|---|
| kimi-k3 | 21.3% | 90.0% | 0.967 | 0.184 |
| gpt-5.6-sol | 21.3% | 100.0% | 0.833 | 0.177 |
| qwen3.8-max | 14.6% | 100.0% | 0.952 | 0.139 |
| glm-5.3 | 21.7% | 90.0% | 0.700 | 0.130 |
| deepseek-v4-pro | 25.0% | 41.7% | 0.931 | 0.097 |
Eight original tasks come from each suite. The author agent writes a new task in that suite's own format; the target model and scoring procedure remain fixed.
Terminal work and software engineering
Official websiteScientific computing
Official websiteBusiness workflows across applications
Official repositoryThe ranking changes across suites. No author agent leads everywhere, so the overall score should not stand in for domain-specific performance.
The benchmark tests whether a delivered task meets acceptance criteria. Whether tasks that pass these checks improve a model after training remains an open question.