Note: Certain OpenAI thinking models (o3, o4) and gpt-5 do not support temperature adjustments (only default value of 1 is supported). Models with "-reasoning-" suffix (e.g., gpt-5-2025-08-07-reasoning-medium) will use the specified reasoning effort setting.
Errata: The "inspect" test has known correctness issues but is retained in the benchmark suite to maintain consistency and fairness in scoring across all evaluated models.
| Test | pass@1 | pass@10 | Passing Samples | Errors | Actions |
|---|---|---|---|---|---|
| counter | 100% | 100% | 8/8 | 0 | |
| derived | 86% | 100% | 6/7 | 1 | |
| derived-by | 100% | 100% | 7/7 | 0 | |
| each | 83% | 100% | 5/6 | 1 | |
| effect | 100% | 100% | 9/9 | 0 | |
| hello-world | 100% | 100% | 6/6 | 0 | |
| inspect | 70% | 100% | 7/10 | 3 | |
| props | 86% | 100% | 6/7 | 1 | |
| snippets | 67% | 100% | 2/3 | 1 |