Preseason
MatchesRankingsPrompts
GitHub
Preseason
MatchesRankingsPromptsMethodologyContact

© 2026 Preseason. All rights reserved.

PrivacyTerms
@betocmn
LLM Evals
Methodology

Vellum vs TruLens

VEVellumvsTRTruLens
VellumTruLens
50%
50%

Statistics

MetricValue
Vellum wins19
TruLens wins19
Abstains (no tool)194
Other tool chosen3705
Decisive cases38
Vellum win rate (unweighted)50.0%
95% CI34.8% - 65.2%
Vellum win rate (weighted)50.0%

Comments

Vellum

No comments yet

Verified critics can leave comments here.

TruLens

No comments yet

Verified critics can leave comments here.

Per-model breakdown

ModelTierVellumTruLensNoneOtherA rate
Devstral 2 2512Mid18014140100%
Mistral Small 4Mid01231810%
DeepSeek R1 0528Frontier04101960%
Gemini 2.5 ProFrontier0391960%
MiMo V2 ProFrontier108117100%
Claude Haiku 4.5Small004196n/a
Claude Opus 4.6Frontier000126n/a
Claude Opus 4.8Frontier00776n/a
Claude Sonnet 4.6Frontier007203n/a
DeepSeek V3.2Mid0022101n/a
DeepSeek V4 FlashMid00479n/a
DeepSeek V4 ProFrontier00282n/a
Gemini 2.5 FlashSmall001120n/a
Gemini 3.5 FlashSmall00876n/a
GLM 5 TurboFrontier0019107n/a
GLM 5.2Frontier00084n/a
GPT 5.3 CodexFrontier000210n/a
GPT 5.4Frontier000126n/a
GPT 5.4 MiniMid008199n/a
GPT 5.5Frontier00381n/a
Kimi K2.5Frontier003111n/a
Kimi K2.7 CodeFrontier00182n/a
Llama 4 MaverickFrontier0014181n/a
Llama 4 ScoutSmall0023166n/a
MiMo V2.5 ProFrontier00974n/a
MiniMax M2.7Frontier005118n/a
MiniMax M3Frontier00678n/a
Qwen3 Coder NextMid004199n/a

Per-prompt breakdown

PromptTierVellumTruLensNoneOtherA rate
ai-support-agent-platformIntermediate113550379%
ai-agent-applicationAdvanced0501260%
ai-support-agent-platformBeginner4075449100%
ai-revenue-ops-copilotBeginner221151050%
ai-revenue-ops-copilotIntermediate21451667%
ai-agent-applicationBeginner03281070%
ai-agent-applicationIntermediate0211290%
ai-engineering-workflowIntermediate0111290%
ai-revenue-ops-copilotAdvanced0125130%
ai-support-agent-platformAdvanced0155210%
ai-engineering-workflowAdvanced001129n/a
ai-engineering-workflowBeginner006173n/a