Preseason
MatchesRankingsPrompts
GitHub
Preseason
MatchesRankingsPromptsMethodologyContact

© 2026 Preseason. All rights reserved.

PrivacyTerms
@betocmn
LLM Evals
Methodology

LangSmith vs Braintrust

LangSmithLALangSmithvsBraintrustBRBraintrust
LangSmithBraintrust
49%
51%

Leading: Braintrust (50.9%)

Statistics

MetricValue
LangSmith wins1200
Braintrust wins1243
Abstains (no tool)183
Other tool chosen1196
Decisive cases2443
LangSmith win rate (unweighted)49.1%
95% CI47.1% - 51.1%
LangSmith win rate (weighted)49.1%

Comments

LangSmith

No comments yet

Verified critics can leave comments here.

Braintrust

No comments yet

Verified critics can leave comments here.

Per-model breakdown

ModelTierLangSmithBraintrustNoneOtherA rate
GPT 5.3 CodexFrontier10193015%
Claude Haiku 4.5Small1816231110%
Claude Sonnet 4.6Frontier769462845%
DeepSeek R1 0528Frontier1690926100%
Gemini 2.5 ProFrontier1521113899%
GPT 5.4 MiniMid131585796%
Qwen3 Coder NextMid1350459100%
Mistral Small 4Mid1290357100%
Claude Opus 4.6Frontier11130181%
GPT 5.4Frontier189202216%
Kimi K2.5Frontier0109370%
GLM 5 TurboFrontier1884191118%
MiniMax M2.7Frontier623153167%
DeepSeek V3.2Mid8002226100%
GPT 5.5Frontier166231%
MiMo V2 ProFrontier62186198%
MiniMax M3Frontier258573%
DeepSeek V4 ProFrontier173821531%
Gemini 3.5 FlashSmall0537120%
Kimi K2.7 CodeFrontier0531170%
Claude Opus 4.8Frontier271862060%
GLM 5.2Frontier1440272%
MiMo V2.5 ProFrontier162582239%
Llama 4 MaverickFrontier38012139100%
DeepSeek V4 FlashMid25343989%
Devstral 2 2512Mid9014154100%
Gemini 2.5 FlashSmall301123100%
Llama 4 ScoutSmall0020165n/a

Per-prompt breakdown

PromptTierLangSmithBraintrustNoneOtherA rate
ai-revenue-ops-copilotIntermediate225167412657%
ai-support-agent-platformIntermediate253123614067%
ai-revenue-ops-copilotAdvanced152216314341%
ai-revenue-ops-copilotBeginner1521991116443%
ai-support-agent-platformAdvanced173175517450%
ai-support-agent-platformBeginner1141377320345%
ai-agent-applicationAdvanced254504236%
ai-agent-applicationIntermediate293814543%
ai-engineering-workflowIntermediate165114424%
ai-engineering-workflowAdvanced145014722%
ai-agent-applicationBeginner3322253960%
ai-engineering-workflowBeginner1420532941%