terminal-bench
View contributors for each Terminal-Bench benchmark.
A benchmark to measure and evolve with the frontier of agent work.
89 high-quality tasks across software engineering, machine learning, security, data science, and more.
The original Terminal-Bench benchmark. 80 tasks testing agents' abilities to complete tasks using a terminal.
A domain-specific benchmark for scientific computing in terminal environments. Currently in development.
Single-task Terminal-Bench challenges spanning inference engine code golf, Rust compiler speedup, and WASM rendering.
A lightweight, diverse, difficult, and high-quality benchmark for agentic evaluation. 82 tasks distilled from 6,627 candidates across 54 benchmarks integrated as Harbor adapters.