Product

Terminal-Bench-Science

Versioned benchmark of 70 scientific research workflows, released August 26, 2026, that scores LLM agents on artifacts produced in a sandboxed terminal.

Coverage

Sep 3, 2026
AINews
Scientists Built a 70-Task Benchmark for AI Agents. Its Launch Leader Solved 30%.

Terminal-Bench-Science measures agents on research workflows, not textbook questions. A vendor-run result released days later shows how quickly its baseline can move.

Sep 3, 2026 · By SCN Staff · 6 min read