July 2026
Turing Terminal-Bench Data
A frontier-calibrated training set of 300 expert-authored, verifier-graded terminal tasks in Machine Learning, Data Analysis, and Scientific Computing, built natively in the Terminal-Bench / Harbor format and validated against the top three frontier models.
48.6% ±1.4
41.3% ±1.4
38.3% ±1.7

