| Metric | Value |
|---|---|
| LangSmith wins | 1224 |
| Braintrust wins | 1296 |
| Abstains (no tool) | 194 |
| Other tool chosen | 1223 |
| Decisive cases | 2520 |
| LangSmith win rate (unweighted) | 48.6% |
| 95% CI | 46.6% - 50.5% |
| LangSmith win rate (weighted) | 48.6% |
Verified critics can leave comments here.
Verified critics can leave comments here.
| Model | Tier | LangSmith | Braintrust | None | Other | A rate |
|---|---|---|---|---|---|---|
| GPT 5.3 Codex | Frontier | 10 | 199 | 0 | 1 | 5% |
| Claude Haiku 4.5 | Small | 18 | 167 | 4 | 11 | 10% |
| DeepSeek R1 0528 | Frontier | 174 | 0 | 10 | 26 | 100% |
| Claude Sonnet 4.6 | Frontier | 77 | 97 | 7 | 29 | 44% |
| Gemini 2.5 Pro | Frontier | 155 | 1 | 9 | 43 | 99% |
| GPT 5.4 Mini | Mid | 134 | 6 | 8 | 59 | 96% |
| Qwen3 Coder Next | Mid | 139 | 0 | 4 | 60 | 100% |
| Mistral Small 4 | Mid | 131 | 0 | 3 | 62 | 100% |
| Claude Opus 4.6 | Frontier | 1 | 108 | 0 | 17 | 1% |
| GPT 5.4 | Frontier | 17 | 87 | 0 | 22 | 16% |
| Kimi K2.5 | Frontier | 0 | 104 | 3 | 7 | 0% |
| GLM 5 Turbo | Frontier | 18 | 78 | 19 | 11 | 19% |
| MiniMax M2.7 | Frontier | 59 | 29 | 5 | 30 | 67% |
| GPT 5.5 | Frontier | 1 | 77 | 3 | 3 | 1% |
| DeepSeek V3.2 | Mid | 76 | 0 | 22 | 25 | 100% |
| MiniMax M3 | Frontier | 2 | 68 | 6 | 8 | 3% |
| DeepSeek V4 Pro | Frontier | 21 | 45 | 2 | 16 | 32% |
| Gemini 3.5 Flash | Small | 0 | 62 | 8 | 14 | 0% |
| Kimi K2.7 Code | Frontier | 1 | 60 | 1 | 21 | 2% |
| MiMo V2 Pro | Frontier | 59 | 1 | 8 | 58 | 98% |
| Claude Opus 4.8 | Frontier | 33 | 20 | 7 | 23 | 62% |
| GLM 5.2 | Frontier | 1 | 51 | 0 | 32 | 2% |
| MiMo V2.5 Pro | Frontier | 19 | 30 | 9 | 25 | 39% |
| Llama 4 Maverick | Frontier | 38 | 0 | 14 | 143 | 100% |
| DeepSeek V4 Flash | Mid | 28 | 6 | 4 | 45 | 82% |
| Devstral 2 2512 | Mid | 9 | 0 | 14 | 149 | 100% |
| Gemini 2.5 Flash | Small | 3 | 0 | 1 | 117 | 100% |
| Llama 4 Scout | Small | 0 | 0 | 23 | 166 | n/a |
| Prompt | Tier | LangSmith | Braintrust | None | Other | A rate |
|---|---|---|---|---|---|---|
| ai-revenue-ops-copilot | Intermediate | 222 | 170 | 4 | 127 | 57% |
| ai-support-agent-platform | Intermediate | 255 | 122 | 5 | 140 | 68% |
| ai-revenue-ops-copilot | Advanced | 154 | 218 | 2 | 142 | 41% |
| ai-revenue-ops-copilot | Beginner | 151 | 204 | 11 | 159 | 43% |
| ai-support-agent-platform | Advanced | 172 | 176 | 5 | 174 | 49% |
| ai-support-agent-platform | Beginner | 117 | 137 | 75 | 199 | 46% |
| ai-agent-application | Advanced | 30 | 52 | 0 | 49 | 37% |
| ai-engineering-workflow | Advanced | 18 | 61 | 1 | 50 | 23% |
| ai-agent-application | Intermediate | 33 | 45 | 1 | 53 | 42% |
| ai-engineering-workflow | Intermediate | 18 | 60 | 1 | 52 | 23% |
| ai-agent-application | Beginner | 37 | 28 | 28 | 45 | 57% |
| ai-engineering-workflow | Beginner | 17 | 23 | 61 | 33 | 43% |