SvelteBench Visualization

Note: OpenAI thinking models (o3, o4) do not support temperature adjustments. o1-pro models use "medium" reasoning effort setting.

← Back to All Results

OpenRouter

deepseek/deepseek-r1-0528

Test Status pass@1 pass@10 Passing Samples Errors Actions
counter ⚠️ PARTIAL 0.3000 1.0000 3/10 22
derived ⚠️ PARTIAL 0.6000 1.0000 6/10 6
derived-by ⚠️ PARTIAL 0.9000 1.0000 9/10 1
each ⚠️ PARTIAL 0.2000 1.0000 2/10 10
effect ✅ PASS 1.0000 1.0000 10/10 0
hello-world ✅ PASS 1.0000 1.0000 10/10 0
inspect ❌ FAIL 0.0000 0.0000 0/10 10
props ⚠️ PARTIAL 0.4000 1.0000 4/10 6
snippets ❌ FAIL 0.0000 0.0000 0/10 10