SvelteBench Visualization

Note: OpenAI thinking models (o3, o4) do not support temperature adjustments. o1-pro models use "medium" reasoning effort setting.

← Back to All Results

OpenRouter

mistralai/devstral-small

Test Status pass@1 pass@10 Passing Samples Errors Actions
counter ⚠️ PARTIAL 0.7000 1.0000 7/10 6
derived ⚠️ PARTIAL 0.1000 1.0000 1/10 13
derived-by ⚠️ PARTIAL 0.1000 1.0000 1/10 13
each ❌ FAIL 0.0000 0.0000 0/10 10
effect ❌ FAIL 0.0000 0.0000 0/10 10
hello-world ⚠️ PARTIAL 0.3000 1.0000 3/10 7
inspect ❌ FAIL 0.0000 0.0000 0/10 13
props ❌ FAIL 0.0000 0.0000 0/10 10
snippets ❌ FAIL 0.0000 0.0000 0/10 10