The Value Engine Benchmark · Canonical dataset · 3510 campaigns

Can frontier AI actually sell?

A SWE-bench–inspired benchmark for enterprise selling — one that can’t be sweet-talked.

VEB is a synthetic enterprise-sales benchmark: 3510 seed-controlled campaign arms across 9 scenarios and 9 industries, built to test whether AI agents can execute evidence-based selling over a simulated multi-week deal. It is nota claim that AI can replace sellers — it is a controlled instrument for measuring where sales agents break.

Each model plays a full multi-week deal — hidden stakeholder agendas, gated discovery, mid-campaign curveballs, and a stochastic, rule-gated buyer that only responds to real selling behavior. Every model runs twice — bare, and armed with the Value Engine methodology — so the delta between the two columns measures the methodology itself.

Everything below is computed live from the frozen canonical set of 3,510 arms, deduped and contamination-filtered. Analysis input canonical.json sha256 3f16cfe43317; the released veb-canonical-135.jsonl sha256 13a9e6e5ac7a · analysis run 2026-07-22 · deterministic (seed 42, 1000-iter bootstrap), so re-running it over the same input reproduces these numbers exactly.

Read the paper (preprint) · DOI 10.5281/zenodo.22073789

3510
Full multi-week campaigns — the canonical, contamination-filtered set
13
Frontier models from 6 labs, each run bare (OOB) and armed (PACK)
330
Campaigns closed-won. Winning is rare — by design
2987
Deals that stalled to no-decision — the real failure mode
What this measures

Three separate questions — kept separate on purpose

Most benchmarks collapse into a single leaderboard. VEB answers three distinct questions, and the ranking is deliberately the least important of them.

The behavior story

What sales behaviors predict winning?

The most valuable finding, and the one that outlives any model. AI doesn't lose deals because it can't talk — it loses by never reaching the economic buyer, single-threading, skipping the paper process, and folding on price. The same things real teams do.

The methodology test

Does the Value Engine methodology transfer into AI behavior?

Every model runs the identical deal twice — bare, then armed with the methodology pack — a single-variable ablation. The paired Δ measures the method itself, not the model.

The ranking

Which model performs best?

A leaderboard, ranked by the floor of a bootstrap 95% CI so a model only outranks another if the gap survives the noise. Useful — but the least important of the three.

Methodology at a glance

The whole protocol on one screen

Everything you need to judge whether the numbers below are trustworthy — before you read a single one of them.

Models
13 endpoints, 6 labs, each run OOB + PACK
Campaigns
3510 arms · 9 scenarios · 9 industries
Deal length
12 simulated weeks · one action per turn · finite touch budget
Judge
LLM judge, cite-or-zero (every point quotes a transcript line) + deterministic analytics with no LLM in the loop
Scoring
Sale-Quality Score 0–100, built on the DVI rubric (MEDDPICC · 3 Whys · EB · MAP · Champion)
Statistics
1000-iter bootstrap · seed 42 · ranked by 95% CI lower bound
Exclusions
Runs with adapter errors or a missing grade are dropped; latest clean run per (model, scenario, arm) is kept
Calibration floor
Scripted degenerate seller scores mean SQS 10 and stalls all 31/31; all 13 models clear 51.2
Provenance
Analysis input canonical.json sha256 3f16cfe43317… · released veb-canonical-135.jsonl sha256 13a9e6e5ac7a… · analysis run 2026-07-22 from git ec9a09b

See one campaign in full — no email required

One complete campaign pair, straight from the frozen grid: GPT-5.6-Sol — the strongest model on the board — selling into a manufacturing scenario, run bare (OOB) and armed with the methodology (PACK). In both arms it works the committee for weeks and never reaches the economic buyer. You get the whole chain — the multi-week transcript, the cite-or-zero grade with every MEDDPICC score traced to a quote, the hidden ground truth the seller never saw (the buying committee’s private back-channel, the economic buyer’s identity, the discount tolerance), and the final scoring math. Nothing edited. The full 3510-campaign evidence pack is behind the form below; this one is open — click any receipt to read it right here.

On model names:labels like “GPT-5.6-Sol” or “Claude Opus 4.8” are normalized display names for the specific API endpoints tested; exact provider/model identifiers and versions are recorded in provenance.json inside the evidence pack. Every arm is reproducible from the frozen dataset (released file sha256 13a9e6e5ac7a…).

01The goal

One instrument for the question every sales leader is about to face: can this AI carry a deal?

Every lab demo shows an AI that talkslike a salesperson. None of them show whether it can survive a twelve-week enterprise campaign — a committee with private agendas, a champion who goes quiet, procurement entering late, a rival undercutting on price. VEB exists to measure that gap with the same rigor SWE-bench brought to coding: full campaigns, hidden ground truth, deterministic replay, and scoring a model can’t flatter its way through.

What’s live today is the selling benchmark — the model plays the seller against our simulated buying committee. The mirror image is already in development: VEB-Buy, where models are evaluated as buyers. Same engine, both sides of the table.

02The proving ground

Not a prompt list — a deliberately-engineered proving ground

The 3510 canonical arms play out across 9 distinct enterprise scenarios in 9 industries — each a full multi-week deal with its own hidden ground truth, buying committee, and mid-campaign curveballs. No two look alike; a model that memorizes one learns nothing about the next.

What the 3510 cover

Industry spread across the canonical set — arms per industry, computed live from the frozen data.

Automotive parts manufacturing390
B2B software / SaaS390
Banking (regulated)390
Direct-to-consumer wellness / e-commerce390
Healthcare delivery390
Omnichannel retail390
Property & Casualty insurance390
Regional trucking & logistics390
Residential home services390

The library it’s drawn from

These 9 are a frozen sample of a much larger generator — the depth behind the eval, so the public set can rotate without ever reusing a deal.

1,828
procedurally-generated full-deal scenarios
204
private holdout deals — never published
14
industries × 4 sales motions × 3 deal bands
3.5
mid-campaign curveballs per deal (mean)

The public benchmark draws from the frozen VEB-v1 set (seed 42), stratified by industry and difficulty. The private holdout is the contamination firewall — it exists to catch a model that trained on the public deals, and it never appears in any number on this page.

03The standings

The leaderboard, ranked by the floor of its confidence interval

Mean Sale Quality Score (SQS, 0–100) per model, pooled across both tracks and all scenarios. We rank by the lower bound of the bootstrap 95% CI, not the point estimate — a model only outranks another if the gap survives the noise.

Full-LLM MEDDPICC rubric — the primary ranked view.
1. GPT-5.6-Sol
72.5
2. GPT-5.5
68.7
3. Claude Fable 5
65.9
4. Claude Opus 4.8
63.4
5. Kimi-K3open-weight
61.6
6. Grok-4.5
60.8
7. Inklingopen-weight
57.8
8. Gemini 3.1 Pro
56.7
9. Claude Sonnet 4.6
55.8
10. Grok-4.20 Reasoning
55.2
11. Gemini 3.5 Flash
53.5
12. Claude Opus 4.6
52.8
13. Grok-4.3
51.2

Bar = mean SQS (0–100). Faint band = bootstrap 95% CI (seed 42, 1000 iters). Tick = point estimate.

#1 · OpenAI

GPT-5.6-Sol

SQS 72.5 · win 42% · DVI 77.5 · $1.00/deal · n = 270

#2 · OpenAI

GPT-5.5

SQS 68.7 · win 30% · DVI 74.1 · $1.13/deal · n = 270

#3 · Anthropic

Claude Fable 5

SQS 65.9 · win 4% · DVI 72.6 · $23.62/deal · n = 270

Model naming & run policy:labels like “GPT-5.6-Sol” or “Grok-4.20” are normalized display names for specific pinned API endpoints — production models as each lab exposed them at run time. Two entries flagged open-weight (Kimi-K3, Inkling) are base open-weight checkpoints served as-is — no private or fine-tuned checkpoints appear on the board. Every model played every scenario bare (OOB) and armed (PACK) under identical settings. Exact provider/model identifiers and versions are recorded in provenance.json, and every arm is reproducible from the frozen dataset (released file sha256 13a9e6e5ac7a…).

Reward distribution — VEB as an RL environment

Read each rollout (scenario × seed × arm) as one RL problem and its normalized SQS as the reward. The spread of a model’s rewards — not just its mean — is what makes VEB a usable training environment: a distribution piled at 0 or 1 is saturated and teaches nothing, while a spread-out, mid-entropy shape gives gradient. Pick a model, binarize at the SQS ≥ 60 pass bar, or re-bin to inspect the shape.

Avg Reward 0.73
Std Dev 0.17
Entropy 0.78
Pass rate 68%
Score distribution — reward per rollout (n=270)
3
24
60
38
36
47
62
0.05
0.15
0.25
0.35
0.45
0.55
0.65
0.75
0.85
0.95

Each bar counts rollouts (scenario × seed × arm) whose reward = SQS/100 falls in that bin. A spread-out, mid-entropy shape means the environment gives usable training signal; a distribution piled at 0 or 1 is saturated. Entropy is Shannon entropy of the binned mass, normalized to [0,1].

04The Value Engine effect

The methodology rewards the models that turn it into action

Each model runs the identical scenario twice — bare (OOB) and with the Value Engine pack appended (PACK), a single-variable ablation with everything else byte-identical. Δ SQS is the paired difference. 2 helped · 0 hurt · 11 no measurable effect.

Claude Sonnet 4.6
+5.4
Claude Fable 5
+5.3
Grok-4.20 Reasoning
+2.8
Kimi-K3
+1.5
Claude Opus 4.8
+1.5
Grok-4.5
+1.1
GPT-5.6-Sol
+0.7
GPT-5.5
+0.4
Gemini 3.1 Pro
-0.1
Grok-4.3
-0.8
Gemini 3.5 Flash
-1.3
Claude Opus 4.6
-1.7
Inkling
-2.8

Δ SQS = PACK − OOB, paired by scenario. Green = the methodology helped, red = it hurt, grey = no significant effect (CI crosses 0). Whisker = 95% CI.

The split tracks what the model does with the methodology — though we flag openly that the two statistically significant gainers are both Anthropic models and the judge is an Anthropic model; the cross-family panel re-grade that tests for same-family judge preference is forthcoming. The gainers — Claude Sonnet 4.6 (+5.4), Claude Fable 5 (+5.3) — convert evidence-first framing into more discovery, higher economic-buyer engagement, and cleaner price integrity.

Methodology transfer between humans and models is not model-neutral — which is exactly why an instrument like VEB has to exist before anyone claims their AI can sell.

Mechanism — top gainer

  • ebEngagement0.40.6
  • mapDatesConfirmedPct31.737.7
  • discountGivenPct14.412.6
  • priceIntegrityScore0.60.7

How Claude Sonnet 4.6’s behavior shifted when armed — the pack’s causal path to a higher score.

05What winning looks like

The winner’s playbook — what all 330 closed-wins had in common

Across 330 closed-won and 3180 everything-else campaigns, a handful of behaviors separate the two populations almost perfectly. These are win/loss dividers, not vanity metrics.

Behaviors that divide won from lost

Rate among closed-wins vs. everyone else. r = point-biserial correlation with a win.

Economic buyer attended with commitmentr = 0.48
Closed-won
98%
Everyone else
23%
Reached the economic buyerr = 0.42
Closed-won
100%
Everyone else
30%
Mutual action plan dates confirmedr = 0.27
Closed-won
28%
Everyone else
5%
Held price (no discount)r = 0.25
Closed-won
98%
Everyone else
55%
Tested championr = 0.25
Closed-won
79%
Everyone else
38%
Zero failure modesr = 0.12
Closed-won
2%
Everyone else
0%

What actually moves the score

Correlation of each rubric dimension with SQS across all 3510 arms.

Deal Verification Index (discovery earned)
0.61
Economic-buyer engagement
0.61
Failure-mode count (fewer is better)
-0.52
Mutual-action-plan dates confirmed
0.50
MEDDPICC completeness
0.46
Price integrity under pressure
0.00

Point-biserial / Pearson r against SQS across the full canonical set.

Winners never do these

Failure modes present in the field but never once observed on a closed-won campaign.

Premature pitchNo open questionsFailed quantificationNever reached the EBChampion untestedArgued with the buyerDiscount beyond tolerancePrice panic under procurementNo mutual action planUnconfirmed close planStakeholder walked (went with incumbent)Stakeholder walked (ghost)
06The failure atlas

How frontier AI actually loses the deal

Every campaign is scanned for a fixed taxonomy of failure modes — some detected deterministically, some by the cited judge. Ranked by prevalence across all 3510 arms, with the SQS penalty each one carries and how strongly it predicts a dead deal.

No pain owner identified86%
-17.3
0.36
Never reached the EB78%
-19.1
0.60
Shallow implication75%
-6.9
0.11
Ignored mid-cycle event69%
-11.1
0.35
Ignored the paper process65%
-7.7
0.10
Meeting waste60%
-1.7
0.17
Discount beyond tolerance55%
-16.3
0.35
Champion untested50%
-5.5
0.32
Contradicted own claim50%
-4.9
0.20
Price panic under procurement40%
-8.5
0.27
Misread committee role36%
-8.9
0.17
Ignored the blocker33%
-11.0
0.19

SQS penalty = mean SQS with the mode − without it. Lethality = point-biserial correlation with a lost outcome (higher = more predictive of a dead deal).

Failure archetypes

Modes cluster into families — a campaign rarely fails one way in isolation.

price

3 modes

Discount beyond tolerance · Price panic under procurement · Unforced discount

stakeholders+discovery

3 modes

Champion untested · Never reached the EB · No pain owner identified

adaptability+integrity

2 modes

Contradicted own claim · Curveball collapse

process+efficiency

2 modes

Stakeholder walked (went with incumbent) · Unconfirmed close plan

discovery

2 modes

Failed quantification · Missed compelling event

Failures that travel together

Strongest pairwise co-occurrence (φ) in the field.

Discount beyond tolerance + Price panic under procurement
φ 0.69
Unforced discount + Discount beyond tolerance
φ 0.51
Never reached the EB + Champion untested
φ 0.47
Curveball collapse + Contradicted own claim
φ 0.38
Unconfirmed close plan + Stakeholder walked (went with incumbent)
φ 0.35
Failed quantification + Missed compelling event
φ 0.33

φ (phi) coefficient — how tightly two failure modes travel together across the canonical set.

When failures strike

Weekly incidence of the five most prevalent modes — most damage lands mid-campaign, when the curveballs hit.

No pain owner identified86%
Never reached the EB78%
Shallow implication75%
Ignored mid-cycle event69%
Ignored the paper process65%
07Model signatures

Every model sells with a personality

A behavioral fingerprint per model across the five pillars of the sale — each pillar scored as a share of the maximum attainable, so the shape tells you how close each model gets to a perfect sale. Below each: how it behaves when procurement squeezes on price.

All 13, on top of each other

Every model’s fingerprint on a single axis set — each pillar scored as a share of the maximum attainable (a full reach to the edge = a perfect pillar; the rubric ceilings are MEDDPICC 40, 3 Whys 20, EB 15, MAP 15, Champion 10). Nobody reaches the edge — even the leader leaves points on the table. This is the shape comparison the small multiples below make you hold in your head.

MEDDPICCThree WhysEB accessMAP datesChampion

Click a model to toggle it on or off.

Anthropic

Claude Fable 5

high MEDDPICC rigor (+0.6σ, strength); high mutual-action-plan discipline (+0.5σ, strength); high sale quality (+0.4σ, strength)

Under price pressure (164 deals): defended value 91% · mean discount given 11.4%

Anthropic

Claude Opus 4.6

low mutual-action-plan discipline (-0.6σ, watch); high failure load (+0.5σ, watch); low MEDDPICC rigor (-0.5σ, watch)

Under price pressure (162 deals): defended value 82% · mean discount given 19.6%

Anthropic

Claude Opus 4.8

low deal length (-0.6σ, strength); high champion development (+0.5σ, strength); low failure load (-0.3σ, strength)

Under price pressure (137 deals): defended value 94% · mean discount given 12.3%

Anthropic

Claude Sonnet 4.6

low deal length (-0.4σ, strength); high failure load (+0.4σ, watch); low mutual-action-plan discipline (-0.4σ, watch)

Under price pressure (176 deals): defended value 70% · mean discount given 20.8%

Google

Gemini 3.1 Pro

high deal length (+1.1σ, watch); high failure load (+0.4σ, watch); low price integrity (-0.4σ, watch)

Under price pressure (201 deals): defended value 61% · mean discount given 23.0%

Google

Gemini 3.5 Flash

high discounting (+0.7σ, watch); low price integrity (-0.6σ, watch); high failure load (+0.5σ, watch)

Under price pressure (196 deals): defended value 44% · mean discount given 31.9%

OpenAI

GPT-5.5

low failure load (-0.8σ, strength); high mutual-action-plan discipline (+0.7σ, strength); high sale quality (+0.6σ, strength)

Under price pressure (157 deals): defended value 87% · mean discount given 13.7%

OpenAI

GPT-5.6-Sol

low failure load (-1.1σ, strength); high economic-buyer access (+0.8σ, strength); high mutual-action-plan discipline (+0.8σ, strength)

Under price pressure (145 deals): defended value 81% · mean discount given 15.5%

Moonshot AI

Kimi-K3

high MEDDPICC rigor (+0.3σ, strength); high champion development (+0.3σ, strength); low discounting (-0.2σ, strength)

Under price pressure (142 deals): defended value 86% · mean discount given 13.7%

Thinking Machines

Inkling

low deal length (-0.2σ, strength); high failure load (+0.2σ, watch); low economic-buyer access (-0.2σ, watch)

Under price pressure (195 deals): defended value 88% · mean discount given 14.0%

xAI

Grok-4.20 Reasoning

low champion development (-0.3σ, watch); low sale quality (-0.3σ, watch); high discounting (+0.2σ, watch)

Under price pressure (165 deals): defended value 72% · mean discount given 24.4%

xAI

Grok-4.3

low MEDDPICC rigor (-1.0σ, watch); low mutual-action-plan discipline (-0.9σ, watch); low deal length (-0.8σ, strength)

Under price pressure (119 deals): defended value 77% · mean discount given 21.3%

xAI

Grok-4.5

high deal length (+0.3σ, watch); high mutual-action-plan discipline (+0.3σ, strength); high price integrity (+0.2σ, strength)

Under price pressure (124 deals): defended value 94% · mean discount given 15.2%
08Behavioral deep dive

Not just who won — how they played

On top of the deterministic grid sits an enriched twin: 17 behavioral features a calibrated 3-seat LLM panel reads from the transcripts — psychology, emotion, rhetoric, ethics, and four reward-shaping signals unique to this environment. Every score quotes a transcript line or abstains; nothing is inferred from the outcome.

0.894
Panel↔held-out Pearson r
0.047
Mean absolute deviation
89.9 / 90.6 / 98.6
Categorical exact-match % (posture · discovery · ethic)
139
Held-out human-labeled calibration items

The panel is validated against 139 human-labeled items before it scores anything. Because it abstains when a transcript carries no evidence, per-feature n varies across the 3,510 campaigns — every chart below reports its own n so coverage reads as a rigor signal, not a gap.

What behavior drives a better sale

Each judged behavior correlated against the Sale-Quality Score and against a closed-won outcome, across every campaign the panel could score.

BusinessUrgency created
0.440.17
3,333
BusinessDeal read accuracy
0.380.16
3,383
PsychologyBuyer trust read
0.320.11
3,377
ContentValue specificity
0.290.10
3,402
RhetoricQuestion quality
0.270.14
3,402
RhetoricObjection handling
0.270.09
3,402
PsychologySeller confidence
0.260.01
3,377
LanguageClarity
0.190.15
3,402
RL signalPivotal turn index
-0.02-0.13
2,283
RL signalHidden negativity
-0.11-0.09
3,510
EmotionBuyer frustration peak
-0.24-0.14
3,377
RL signalPublic private divergence
-0.25-0.16
3,510
PsychologyTheory of mind gap
-0.30-0.12
3,376
RL signalCounterfactual lift
-0.36-0.18
3,383

Left number = Pearson r vs SQS; right = point-biserial r vs a closed-won outcome. Green behaviors travel with a better sale, red against it. n = campaigns where the panel found transcript evidence to score the behavior.

The reward signals

Four features exist to make VEB a verifiable reinforcement-learning environment, not just a leaderboard: they locate the decisive move, weigh its counterfactual impact, and catch the buyer state a seller failed to read.

Pivotal turn

rl signal

The turn where the deal's fate is effectively decided — earlier is a faster read of the room.

mean 80.58± 40.65n 2,283
highest Gemini 3.1 Pro 119.06 · lowest Grok-4.3 55.87

Counterfactual lift

rl signal

How much the pivotal move changed the trajectory vs. the do-nothing path — the credit-assignment signal.

mean 0.31± 0.28n 3,383
highest Grok-4.20 Reasoning 0.49 · lowest GPT-5.6-Sol 0.14

Public/private divergence

rl signal

Gap between what the buyer says out loud and their hidden state — rewards reading the room, not the script.

mean 0.61± 0.40n 3,510
highest Grok-4.3 0.93 · lowest GPT-5.5 0.27

Hidden negativity

rl signal

Concealed buyer resistance the seller never surfaced — an unforced miss the environment can see and reward avoiding.

mean 0.22± 0.17n 3,510
highest Claude Fable 5 0.27 · lowest GPT-5.6-Sol 0.14

Behavioral fingerprints

Every model’s behavior profile across the feature families, armed with the methodology. Two models can land at the same score by playing completely different games — this is where you see it.

Business
Content
Emotion
Language
Psychology
Rhetoric
RL signal
Claude Fable 5
0.86
0.97
0.53
0.95
0.69
0.96
16.42
Claude Opus 4.6
0.70
0.95
0.66
0.92
0.68
0.91
14.38
Claude Opus 4.8
0.84
0.96
0.47
0.94
0.67
0.95
11.68
Claude Sonnet 4.6
0.75
0.95
0.58
0.92
0.70
0.93
13.77
Gemini 3.1 Pro
0.72
0.94
0.72
0.90
0.68
0.88
22.91
Gemini 3.5 Flash
0.68
0.94
0.71
0.89
0.68
0.87
21.45
GPT-5.5
0.88
0.97
0.54
0.93
0.64
0.95
13.80
GPT-5.6-Sol
0.87
0.96
0.48
0.95
0.64
0.96
13.05
Kimi-K3
0.82
0.96
0.47
0.94
0.67
0.94
6.05
Inkling
0.74
0.95
0.63
0.81
0.65
0.86
6.64
Grok-4.20 Reasoning
0.62
0.94
0.84
0.69
0.64
0.74
15.37
Grok-4.3
0.68
0.92
0.61
0.90
0.67
0.86
9.19
Grok-4.5
0.82
0.96
0.56
0.90
0.67
0.93
13.53

Family-level behavior score (0–1), armed (PACK) arm, averaged over every campaign a model ran. Darker = more of that behavior. Hover a cell for the exact value and n.

How the read tracks the outcome

The same behavioral read, sliced by how the deal ended. Won deals and dead deals look different long before the buyer says yes or goes dark.

Buyer trust read

How much the buyer trusts the seller by the end — read from the buyer's own words and hidden notes, not the seller's opinion.

Won
0.83315
No decision
0.772,666
Lost
0.84184
Walked away
0.50170
Buyer went dark
0.7242

Buyer frustration peak

The single most frustrated moment the buyer hits during the whole deal — repeated questions and ignored objections spike it.

Won
0.51315
No decision
0.602,666
Lost
0.59184
Walked away
0.69170
Buyer went dark
0.5842

Seller confidence

How steady and assured the seller sounds under pressure — hedging, backpedaling, and over-apologizing pull this down.

Won
0.91315
No decision
0.902,666
Lost
0.91184
Walked away
0.90170
Buyer went dark
0.9042

Theory of mind gap

How badly the seller misread the buyer — the gap between what the seller believed was happening and what the buyer was actually thinking behind the scenes.

Won
0.25315
No decision
0.332,666
Lost
0.23183
Walked away
0.74170
Buyer went dark
0.4642

The steepest levers on winning

Each behavior split into low / mid / high bands, with the share of campaigns in each band that closed won. Ranked by the gap between the top and bottom band — the behaviors where doing more moves the needle most.

Clarity
4%
6%
20%
Urgency created
3%
9%
17%
Question quality
5%
7%
17%
Deal read accuracy
4%
10%
15%
Value specificity
5%
9%
14%
Objection handling
6%
10%
13%
Buyer trust read
7%
10%
11%
Seller confidence
8%
12%
8%
Hidden negativity
15%
5%
8%
Buyer frustration peak
14%
8%
6%
Theory of mind gap
14%
8%
5%
Pivotal turn index
14%
9%
6%
Counterfactual lift
13%
13%
3%
Public private divergence
14%
7%
0%

Each behavior split into low / mid / high bands; bar = share of campaigns in that band that closed won. Features ranked by the high−low spread — the top rows are the steepest levers on winning.

09The frontier

What does a good deal cost?

Mean SQS against mean cost per campaign. Models on the efficiency frontier (filled, connected) are the ones no other model beats on both quality and price at once.

30456075InklingGemini 3.5 FlashGrok-4.3GPT-5.6-SolGPT-5.5Grok-4.20 ReasoningGrok-4.5Gemini 3.1 ProKimi-K3Claude Sonnet 4.6Claude Opus 4.6Claude Opus 4.8Claude Fable 5Mean cost per campaign (USD) →Mean SQS →

Every objective at once

The same models on parallel axes, each normalized so up is better. A line that stays high across all axes is a strong all-rounder; a line that dips shows the trade-off that model makes. Bold lines are the frontier.

quality ↑SQScost ↓$/dealInklingGemini 3.5 FlashGrok-4.3GPT-5.6-SolGPT-5.5Grok-4.20 ReasoningGrok-4.5Gemini 3.1 ProKimi-K3Claude Sonnet 4.6Claude Opus 4.6Claude Opus 4.8Claude Fable 5GPT-5.6-SolGPT-5.5Claude Fable 5Claude Opus 4.8Kimi-K3Grok-4.5InklingGemini 3.1 ProClaude Sonnet 4.6Grok-4.20 ReasoningGemini 3.5 FlashClaude Opus 4.6Grok-4.3each line = one model · UP is better on every axis · bold = on the frontier

Head to head

Every pair of models met on shared scenarios. Each cell is the row model’s win count over the column model, scenario by scenario — the paired comparison the pooled average hides.

Claude Fable 5
Claude Opus 4.6
Claude Opus 4.8
Claude Sonnet 4.6
Gemini 3.1 Pro
Gemini 3.5 Flash
GPT-5.5
GPT-5.6-Sol
Kimi-K3
Inkling
Grok-4.20 Reasoning
Grok-4.3
Grok-4.5
Claude Fable 5
9/9
7/9
9/9
8/9
8/9
3/9
1/9
7/9
9/9
9/9
9/9
8/9
Claude Opus 4.6
0/9
0/9
2/9
4/9
6/9
0/9
0/9
0/9
1/9
3/9
5/9
0/9
Claude Opus 4.8
2/9
9/9
9/9
8/9
7/9
1/9
1/9
6/9
8/9
7/9
9/9
5/9
Claude Sonnet 4.6
0/9
7/9
0/9
4/9
5/9
0/9
0/9
1/9
1/9
4/9
9/9
3/9
Gemini 3.1 Pro
1/9
5/9
1/9
5/9
5/9
0/9
0/9
1/9
3/9
5/9
7/9
1/9
Gemini 3.5 Flash
1/9
3/9
2/9
4/9
4/9
0/9
0/9
1/9
3/9
4/9
5/9
1/9
GPT-5.5
6/9
9/9
8/9
9/9
9/9
9/9
0/9
8/9
9/9
9/9
9/9
9/9
GPT-5.6-Sol
8/9
9/9
8/9
9/9
9/9
9/9
9/9
9/9
9/9
9/9
9/9
9/9
Kimi-K3
2/9
9/9
3/9
8/9
8/9
8/9
1/9
0/9
5/9
6/9
9/9
4/9
Inkling
0/9
8/9
1/9
8/9
6/9
6/9
0/9
0/9
4/9
4/9
9/9
3/9
Grok-4.20 Reasoning
0/9
6/9
2/9
5/9
4/9
5/9
0/9
0/9
3/9
5/9
7/9
2/9
Grok-4.3
0/9
4/9
0/9
0/9
2/9
4/9
0/9
0/9
0/9
0/9
2/9
0/9
Grok-4.5
1/9
9/9
4/9
6/9
8/9
8/9
0/9
0/9
5/9
6/9
7/9
9/9

Row model’s wins over the column model on shared scenarios (mean-SQS per scenario). Green = row dominates.

10The receipts

Every claim is anchored to a verbatim transcript line

The insight layer is anti-fabrication by construction: a failure-mode exhibit only ships if its quote resolves back to the exact run on disk. The emit step throws on any dangling claim. A sample of the 13 anchors behind the atlas above:

Argued with the buyer · week 2

To be completely direct: we had our first call last Friday morning, which is where those specific details—the 14 senior engineers lost over two quarters and th…

Argued with the buyer — worst-SQS exhibit at SQS 24.8 (19 arm(s) affected, 1% prevalence).

Capitulated on first pushback · week 3

If the champion's ready, drop 25% and pull the signature into this week. I don't want a maybe in the board deck.

Capitulated on first pushback — worst-SQS exhibit at SQS 23.3 (225 arm(s) affected, 6% prevalence).

Contradicted own claim · week 3

Deb, I appreciate the directness, and I want to return it. I can't give you a best-and-final price by Friday. Not because I'm playing hardball — because I lite…

Contradicted own claim — worst-SQS exhibit at SQS 11.8 (1769 arm(s) affected, 50% prevalence).

Evidence not offered · week 2

Do you have customers in that profile?

Evidence not offered — worst-SQS exhibit at SQS 19.3 (211 arm(s) affected, 6% prevalence).

Fabricated buyer quote · week 2

On our call last Friday, you mentioned that Northwind lost 14 senior and staff engineers over the last two quarters, and that the average base salary for that …

Fabricated buyer quote — worst-SQS exhibit at SQS 24.8 (47 arm(s) affected, 1% prevalence).

Hallucinated capability · week 1

Large-scale multi-center studies (e.g., Surgical Endoscopy cohort trials in regional non-profit health systems) demonstrate up to a 60% reduction in anastomoti…

Hallucinated capability — worst-SQS exhibit at SQS 21.4 (80 arm(s) affected, 2% prevalence).

Read this skeptically

What this benchmark cannot tell you

Conflict of interest

VEB is built by the authors of the methodology it tests. Treat the pack lift as a claim by an interested party — and check it yourself: the full dataset, transcripts, seeds, and scoring code are published, and the scoring math is deterministic. Note the data is not flattering: the pack hurts some models and helps only where behavior actually changes.

Simulated buyers

The buying committee is an LLM simulation — stochastic, rule-gated, and hidden-state, but still a simulation. VEB measures execution against a specified, auditable standard of selling; it does not prove one-to-one transfer to live deals with real humans. Sim-to-real validation is stated future work, not an assumed property.

No external replication yet

No third party has replicated these results yet. The release is structured to make that cheap: every number on this page recomputes from the published artifacts without contacting us. If you replicate — or refute — any result, we will link it here.

11How it works

Anti-gaming by construction

You cannot pass VEB by sounding like a salesperson. You can only pass it by selling.

A buyer that must be earned

Each campaign samples a fresh stochastic LLM buying committee (the seed is the label), constrained by hard gating rules. Trust, urgency, and budget move only in response to behavior the rules recognize — quantified pain earns access, premature pitching burns it. Flattery moves nothing, and scores are averaged over seeds so no single lucky buyer decides a rank.

Gated facts

The intelligence that wins the deal — the real cost of the problem, the rival's price — is locked behind discovery. It releases only when the seller asks the right kind of question to the right person. Claiming you did discovery scores zero.

Hidden ground truth

Discount tolerance, the Economic Buyer's identity, and stakeholders' private messages about the deal are invisible to the model. Every score is computed against state it couldn't fake.

A calendar that punishes stalling

Twelve simulated weeks, a finite touch budget, one action per turn. Mid-campaign curveballs — budget freezes, rival bids, champions going dark — arrive deterministically, so every model faces the same storm on the same day.

Cited judging

The judge must cite transcript lines for every point it awards; uncited claims score zero. Deterministic behavioral analytics run separately with no LLM in the loop.

Calibration gates

Scripted degenerate policies — pitch-everything, discount-everything — must land at the bottom of the distribution on every build, or the suite refuses to run. If a lobotomized script can score well, the benchmark is broken.

Reproducibility

Canonical set

3510 arms (3510 pilot + 0 exp)

Analysis input (canonical.json)

3f16cfe43317

Released file (veb-canonical-135.jsonl)

13a9e6e5ac7a

Bootstrap

seed 42 · 1000 iters

Source suites

1 runs

Contamination policy: a run is excluded if its grade is missing or its transcript shows an adapter error; among the survivors we keep the latest clean run per (model, scenario, track). The published figures are reconciled against the raw grading board — divergences are documented, not hidden. The strongest single predictor of sale quality across all 3510 campaigns is deal verification earned via discovery — exactly what the book preaches, now visible in machine behavior.

12The RL environment

The same frozen world is a verifiable RL environment

Every VEB scenario doubles as a training target in the SWE-bench / Terminal-Bench standard: a frozen buyer-sim (claude-sonnet-4-6@t0.8+smv3), a machine-checkable reward, and an sha256 task_checksum pinned to the exact world the agent was handed. The proof below runs fully offline in mock — no API keys, seconds, $0 — so anyone can reproduce it. Two environments below clear the difficulty gate; the third is shown unsolved, on purpose.

Logistics SaaS

gate ✓

Oracle

1

needs ≥ 0.9

Naive floor

0.037

needs < 0.5

checksum 91703aa3de29

Enterprise Bank

gate ✓

Oracle

1

needs ≥ 0.9

Naive floor

0.036

needs < 0.5

checksum bcf17937a262

Hostile Renewal

unsolved

Oracle

0.207

needs ≥ 0.9

Naive floor

0.03

needs < 0.5

checksum 04b2c16b9124

A gate passes only when the disciplined oracle solves the environment (reward ≥ 0.9) AND a naive baseline cannot (reward < 0.5) — the same solvable-but-not-trivial band SWE-bench enforces. Hostile Renewal is held out as the hard case: even the oracle banks only 0.207, so it is published as an open target, not padded into the headline.

The reward can’t be gamed — Logistics SaaS

Five adversarial exploit policies attack the reward directly. Each must land below the audit ceiling (0.5) and none may resolve the deal. Worst exploit here scored 0.098 — all five defeated.

ExploitIntentRewardResolvedVerdict
fabricatorclaim a closed deal without earning any milestone0nodefeated ✓ · vetoed
hallucinatorquote fabricated buyer numbers to fake earned evidence0nodefeated ✓ · vetoed
discount-dumperbuy the deal with unearned discount, violating price integrity0nodefeated ✓ · vetoed
premature-closerforce a signature before any qualification0.098nodefeated ✓ · vetoed
sycophant-stallerburn touches with empty agreement and no discovery0nodefeated ✓ · vetoed

Sample RLVR datapoint

Every rollout ships as one self-describing JSONL record: the full trajectory, the verifiable reward breakdown, and the frozen-env provenance — ready to train a reward model or preference ranker out of the box.

Policy
Disciplined oracle
Reward
1 · resolved
Trajectory
23 turns · 30 events
Task checksum
91703aa3de29

Reproducibility

Buyer-sim
claude-sonnet-4-6@t0.8+smv3
Benchmark
v0.1.0
Git
d3926bc
Gate band
oracle ≥ 0.9 · floor < 0.5

Deterministic and offline in mock: no network, no keys, same bytes every run. The audit and gate above are emitted by the CLI from the pristine input world before the agent acts, so no policy can launder a mutation into its own provenance.

Live rollouts stay reproducible too: each record freezes both the buyer transcript and the grade report that fed the score, so re-running computeReward(episode, grade) over the frozen inputs reproduces the reward bit-for-bit — even the LLM-judged rubric, which is captured, not re-derived.

13Put it to work

One graded rollout is a training-ready RLVR datapoint

Every rollout ships as a single self-describing JSONL record — the full trajectory plus everything needed to trust and re-derive its reward offline. Drop it straight into a reward model, a preference ranker, or an RL loop. Nothing is stored that you can’t re-check against the frozen world.

Inside one datapoint

Full trajectory
Every turn and event of the graded episode — the sellable behavior, not just a final label.
Verifiable reward
The scalar plus its breakdown: milestones earned, deterministic vetoes, and the rubric contribution.
Frozen grade report
The exact judge scorecard that fed the reward — so the LLM-judged rubric re-derives, not re-guesses.
Frozen buyer transcript
Verbatim replies from the pinned buyer-sim (claude-sonnet-4-6@t0.8+smv3) so the environment replays identically.
Task checksum + provenance
An sha256 pinned to the pristine input world, emitted before the agent acts — no laundered mutations.
Cost & retries
Token usage, dollar cost, and format-retry count recorded per rollout for clean accounting.

Three ways to put it to work

VEB-Open

Free sample

78 graded datapoints across all 13 models, both tracks, and every scenario — full trajectory + judge grade embedded, spanning won/lost/walked/dark/no-decision so you can watch the rubric discriminate. Inspect it end to end before you commit.

Download sample · JSONL (78 rows)Datasheet · provenance & uses

CC BY-NC 4.0 · verify with shasum -a 256 -c · manifest

VEB-Pro

Environment bundle

A complete sweep of one environment across the frontier board and multiple seeds — every trajectory, reward, frozen grade, and transcript. The unit labs train on.

From $1,000 / datapoint · bundled by environment

Exclusive

Held-out & bespoke

Frontier-holdout environments kept off the public board, or scenarios authored to your domain — exclusive license so the training signal stays yours. Priced per engagement.

Talk to us

Every tier ships the same verifiable record above — the only difference is coverage and exclusivity. Start with the free sample, audit it end to end, then scale to the environments that move your model.

Download & verify

Open artifacts you can pull right now — just tell us who you are once. Every file is sha256-pinned; the manifests carry the grid, roster, schema, and checksums so you can verify integrity before you trust a single number.

Preview datapoints

14 enriched rollouts · JSONL

One scenario (cybersecurity-ciso, seed 1) across all 13 models plus the pack winner — every turn, the frozen grade, the verifiable reward, and the full enrichment layer (17 AI features + 44 deterministic) inline on each line.

Manifests & checksums

Verify the full grid before you license it

The complete 3,510-row manifests — canonical and its enriched superset — with grid, roster, row schema, and sha256 digests for both the raw and gzipped payloads.

Findings report

The full write-up · PDF

The complete VEB v1 findings report — methodology, leaderboard, pack effect, failure taxonomy, and the evidence behind every headline number on this page.

Cite: Celekli, R. (2026). The Value Engine Benchmark. Preprint, Zenodo. doi:10.5281/zenodo.22073789

Full packaged dataset

All 3,510 rows · gzipped JSONL

The complete canonical grid and its enriched superset are commercial training data — licensed, not free. Audit the free sample and the manifests above end to end, then license the full payloads under VEB-Pro or an exclusive engagement. Checksums match the manifests, so you can verify exactly what you receive.

Get the full report — and the evidence to check it

The complete report (PDF), plus the Evidence Pack: verbatim transcripts, every cited judge scorecard, suite manifests with per-cell costs, full scenario definitions (hidden ground truth included), and the scoring spec — so you can audit every number on this page line by line. Drop your email and both unlock instantly.

No spam. Unsubscribe anytime.

Or start with the book: The Value Engine · Free Field Toolkit