Distinct sequence-first benchmark construction
TASTE samples and selects tool sequences before instantiating scenarios, making procedural coverage an explicit construction target rather than a post hoc description.
↳ Introduction; §§2.1 and 3
Assembling the evidence…
As agent capabilities advance, existing benchmarks, such as $τ^2$-Bench, are becoming increasingly saturated. Yet constructing new benchmark tasks remains complex, costly, and labor-intensive. Moreover, the standard approach, in which scenarios are first written in natural language and then mapped to tool sequences, captures only a narrow subset of the tool-use patterns agents exercise. In this paper, we address these problems by reversing the task construction process. We propose TASTE: Task Synthesis from Tool Sequence Evolution, an automatic method that generates challenging tasks with broader tool-use coverage. TASTE utilizes an Adaptive Contrastive $n$-gram model trained on LLM-judged validity signals. This enables sampling valid tool sequences that cover a vast range of tool combinations. TASTE then selects representative sequences from the pool via clustering, instantiates them into complete benchmark tasks, and refines them through iterative difficulty evolution. Using TASTE, we construct $τ^c$-Bench, a challenging extension of the three domains of $τ^2$-Bench. We evaluate $11$ agent/user LLM pairs and find that models nearly saturating $τ^2$-Bench suffer severe performance drops on our tasks (e.g., Gemini-3-Flash falls from $0.82\!-\!0.94$ to $0.28\!-\!0.61$). Beyond increasing difficulty, our generated tasks more than double the number of unique tool combinations agents must execute. Our results suggest high scores on existing benchmarks often reflect saturation rather than robust task-solving ability. By automating the generation of difficult, high-coverage benchmarks, TASTE enables continuous, scalable evaluation of future agents.
models nearly saturating τ²-Bench suffer severe performance drops on our tasks
observed drops are clear, but longer and adversarial tasks prevent uniquely attributing them to saturation
our generated tasks more than double the number of unique tool combinations agents must execute
sequence diagnostics show large diversity gains, though magnitudes vary by domain and metric
high scores on existing benchmarks often reflect saturation rather than robust task-solving ability
the cross-benchmark design cannot separate saturation from intentional changes in task length and conversational adversity
We achieved 86.7% validity rate with our full trained model compared to uniform tool sampling (6.7%)
the reported Airline ablation directly compares uniform sampling with the full adaptive contrastive model
Derived from the full evaluation — not a separate score.
Strengths
TASTE samples and selects tool sequences before instantiating scenarios, making procedural coverage an explicit construction target rather than a post hoc description.
↳ Introduction; §§2.1 and 3
The paper supplies pseudocode, exact prompts, consolidated hyperparameters, API identifiers, task counts, and generation and evaluation costs.
↳ §4; Tables 3–4; Appendices B and D
Coverage is evaluated with weighted edit distance, type-token ratios, entropy, and successful simulation trajectories, while the verifier receives a separate precision-recall assessment and all 15 universally failed tasks are manually checked.
↳ §§5–5.1; Figures 2–3; Appendix A, Table 2
Limitations
τᶜ-Bench tasks are much longer and are deliberately evolved with decoys, information withholding, and policy pressure. The comparison therefore establishes lower performance on a harder distribution but does not uniquely identify prior-benchmark saturation as the cause.
↳ §3.3; Table 1; Appendix A, Table 2
Table 1 primarily reports single-trial pass¹ point estimates without confidence intervals, repeated-run uncertainty, or statistical tests; pass³ is limited to Airline.
↳ §4; Table 1; NeurIPS Checklist Q7
Airline, Retail, and Telecom all come from the same conversational customer-service benchmark family. The claimed adaptability to single-turn and non-conversational settings is not empirically demonstrated.
↳ §4; §7
The paper's strongest contribution is the sequence-first construction framework and its explicit treatment of coverage through tool-sequence diagnostics. Sections 3–5 and the appendices provide enough implementation detail to make the benchmark artifact unusually transparent, while Figure 3 supplies useful but incomplete component analyses. The principal interpretive problem is that τᶜ-Bench changes sequence length and conversational adversity alongside coverage, so Table 1 cannot by itself establish that lower scores specifically diagnose saturation. This central limitation, together with single-trial point estimates and evidence confined to one benchmark family, keeps the overall quality in the gap-bearing rather than fully well-supported range.
Nabu’s assessment, alongside the field’s view.
Are you an author of this paper?
Sound3.3
Confidence highThe score reflects a distinct sequence-first benchmark-construction method, an operational definition of coverage, and a new three-domain benchmark. Its broader conceptual reach is limited by evidence drawn only from the inherited τ²-Bench environments.
We sample a diverse set of tool sequences and synthesize the full tasks around them.
The pipeline, prompts, hyperparameters, model identifiers, costs, and validation checks are extensively documented. The score is reduced for incomplete end-to-end component ablations and performance estimates reported without confidence intervals or repeated-run uncertainty.
We did not run additional seeds per task pair due to API-cost constraints.
The paper gives a coherent progression from desiderata through pipeline stages to evaluation. The saturation interpretation is stated more strongly than the comparison identifies because the new tasks also differ substantially in length and adversarial design.
high scores on existing benchmarks often reflect saturation rather than robust task-solving ability
Related work distinguishes TASTE from ToolGrad, Trajectory2Task, and the τ-Bench lineage. The paper does not trace how deliberate changes in task length and conversational adversity constrain attribution of lower scores to benchmark saturation, despite reporting those changes.
The generated base tasks do not reflect realistic conditions
Lower confidence on Methodological Rigour, Positioning — domain match limited.
Caveats4 of 4 checks
The reported values are internally traceable, but two headline summaries combine or generalize across heterogeneous cells. These presentation choices do not overturn the direction of the results.
The study uses synthetic benchmark environments, declares LLM use and funding, and claims release of code and generated data. No human-subject or crowdsourcing procedures requiring ethics approval are reported.
Flags: 3 declared / 5 total
36 of 43 checkable references verified
54 references in manuscript 9 are books, websites or datasets — counted, but not index-checkable 2 could not be checked — not the same as no match found
No retraction notice found in Retraction Watch.
Sources: Retraction Watch ✓
Where this paper’s evidence sits on the path from initial observation to real-world use.
The artifact is evaluated with 11 agent-user pairs in the established τ²-Bench harness, with code, data, prompts, costs, and model identifiers reported. It remains a simulated benchmark rather than a live operational deployment.
AI-generated, human-governed. Something look off? Contact us to request a review.