Model Library / Benchmarks / τ²-bench (Telecom)

Coding & Agents

τ²-bench (Telecom)

Sierra's dual-control benchmark for conversational tool-using agents that must follow domain policy while guiding a user who also acts in the shared environment. Source: τ²-Bench paper (arXiv:2506.07982), Telecom domain.

Features measured

Tool use Reasoning

Use cases unlocked

Tool use & function calling Long-horizon agents Everyday agent tasks

Leaderboard

Higher is better · pass^1 %
1 Claude 3.7 Sonnet 49%
2 GPT-4.1 mini 44%
3 o4-mini 42%
4 GPT-4.1 34%

Benchmark source ↗

← All benchmarks