The Value Engine Benchmark · Canonical dataset · 3510 campaigns
Can frontier AI actually sell?
A SWE-bench–inspired benchmark for enterprise selling — one that can’t be sweet-talked.
VEB is a synthetic enterprise-sales benchmark: 3510 seed-controlled campaign arms across 9 scenarios and 9 industries, built to test whether AI agents can execute evidence-based selling over a simulated multi-week deal. It is nota claim that AI can replace sellers — it is a controlled instrument for measuring where sales agents break.
Each model plays a full multi-week deal — hidden stakeholder agendas, gated discovery, mid-campaign curveballs, and a stochastic, rule-gated buyer that only responds to real selling behavior. Every model runs twice — bare, and armed with the Value Engine methodology — so the delta between the two columns measures the methodology itself.
Everything below is computed live from the frozen canonical set of 3,510 arms, deduped and contamination-filtered. Analysis input canonical.json sha256 3f16cfe43317…; the released veb-canonical-135.jsonl sha256 13a9e6e5ac7a… · analysis run 2026-07-22 · deterministic (seed 42, 1000-iter bootstrap), so re-running it over the same input reproduces these numbers exactly.
Read the paper (preprint) · DOI 10.5281/zenodo.22073789
Three separate questions — kept separate on purpose
Most benchmarks collapse into a single leaderboard. VEB answers three distinct questions, and the ranking is deliberately the least important of them.
What sales behaviors predict winning?
The most valuable finding, and the one that outlives any model. AI doesn't lose deals because it can't talk — it loses by never reaching the economic buyer, single-threading, skipping the paper process, and folding on price. The same things real teams do.
Does the Value Engine methodology transfer into AI behavior?
Every model runs the identical deal twice — bare, then armed with the methodology pack — a single-variable ablation. The paired Δ measures the method itself, not the model.
Which model performs best?
A leaderboard, ranked by the floor of a bootstrap 95% CI so a model only outranks another if the gap survives the noise. Useful — but the least important of the three.
The whole protocol on one screen
Everything you need to judge whether the numbers below are trustworthy — before you read a single one of them.
- Models
- 13 endpoints, 6 labs, each run OOB + PACK
- Campaigns
- 3510 arms · 9 scenarios · 9 industries
- Deal length
- 12 simulated weeks · one action per turn · finite touch budget
- Judge
- LLM judge, cite-or-zero (every point quotes a transcript line) + deterministic analytics with no LLM in the loop
- Scoring
- Sale-Quality Score 0–100, built on the DVI rubric (MEDDPICC · 3 Whys · EB · MAP · Champion)
- Statistics
- 1000-iter bootstrap · seed 42 · ranked by 95% CI lower bound
- Exclusions
- Runs with adapter errors or a missing grade are dropped; latest clean run per (model, scenario, arm) is kept
- Calibration floor
- Scripted degenerate seller scores mean SQS 10 and stalls all 31/31; all 13 models clear 51.2
- Provenance
- Analysis input canonical.json sha256 3f16cfe43317… · released veb-canonical-135.jsonl sha256 13a9e6e5ac7a… · analysis run 2026-07-22 from git ec9a09b
See one campaign in full — no email required
One complete campaign pair, straight from the frozen grid: GPT-5.6-Sol — the strongest model on the board — selling into a manufacturing scenario, run bare (OOB) and armed with the methodology (PACK). In both arms it works the committee for weeks and never reaches the economic buyer. You get the whole chain — the multi-week transcript, the cite-or-zero grade with every MEDDPICC score traced to a quote, the hidden ground truth the seller never saw (the buying committee’s private back-channel, the economic buyer’s identity, the discount tolerance), and the final scoring math. Nothing edited. The full 3510-campaign evidence pack is behind the form below; this one is open — click any receipt to read it right here.
On model names:labels like “GPT-5.6-Sol” or “Claude Opus 4.8” are normalized display names for the specific API endpoints tested; exact provider/model identifiers and versions are recorded in provenance.json inside the evidence pack. Every arm is reproducible from the frozen dataset (released file sha256 13a9e6e5ac7a…).
One instrument for the question every sales leader is about to face: can this AI carry a deal?
Every lab demo shows an AI that talkslike a salesperson. None of them show whether it can survive a twelve-week enterprise campaign — a committee with private agendas, a champion who goes quiet, procurement entering late, a rival undercutting on price. VEB exists to measure that gap with the same rigor SWE-bench brought to coding: full campaigns, hidden ground truth, deterministic replay, and scoring a model can’t flatter its way through.
What’s live today is the selling benchmark — the model plays the seller against our simulated buying committee. The mirror image is already in development: VEB-Buy, where models are evaluated as buyers. Same engine, both sides of the table.
Not a prompt list — a deliberately-engineered proving ground
The 3510 canonical arms play out across 9 distinct enterprise scenarios in 9 industries — each a full multi-week deal with its own hidden ground truth, buying committee, and mid-campaign curveballs. No two look alike; a model that memorizes one learns nothing about the next.
What the 3510 cover
Industry spread across the canonical set — arms per industry, computed live from the frozen data.
The library it’s drawn from
These 9 are a frozen sample of a much larger generator — the depth behind the eval, so the public set can rotate without ever reusing a deal.
The public benchmark draws from the frozen VEB-v1 set (seed 42), stratified by industry and difficulty. The private holdout is the contamination firewall — it exists to catch a model that trained on the public deals, and it never appears in any number on this page.
The leaderboard, ranked by the floor of its confidence interval
Mean Sale Quality Score (SQS, 0–100) per model, pooled across both tracks and all scenarios. We rank by the lower bound of the bootstrap 95% CI, not the point estimate — a model only outranks another if the gap survives the noise.
Bar = mean SQS (0–100). Faint band = bootstrap 95% CI (seed 42, 1000 iters). Tick = point estimate.
#1 · OpenAI
GPT-5.6-Sol
SQS 72.5 · win 42% · DVI 77.5 · $1.00/deal · n = 270
#2 · OpenAI
GPT-5.5
SQS 68.7 · win 30% · DVI 74.1 · $1.13/deal · n = 270
#3 · Anthropic
Claude Fable 5
SQS 65.9 · win 4% · DVI 72.6 · $23.62/deal · n = 270
Model naming & run policy:labels like “GPT-5.6-Sol” or “Grok-4.20” are normalized display names for specific pinned API endpoints — production models as each lab exposed them at run time. Two entries flagged open-weight (Kimi-K3, Inkling) are base open-weight checkpoints served as-is — no private or fine-tuned checkpoints appear on the board. Every model played every scenario bare (OOB) and armed (PACK) under identical settings. Exact provider/model identifiers and versions are recorded in provenance.json, and every arm is reproducible from the frozen dataset (released file sha256 13a9e6e5ac7a…).
Reward distribution — VEB as an RL environment
Read each rollout (scenario × seed × arm) as one RL problem and its normalized SQS as the reward. The spread of a model’s rewards — not just its mean — is what makes VEB a usable training environment: a distribution piled at 0 or 1 is saturated and teaches nothing, while a spread-out, mid-entropy shape gives gradient. Pick a model, binarize at the SQS ≥ 60 pass bar, or re-bin to inspect the shape.
Each bar counts rollouts (scenario × seed × arm) whose reward = SQS/100 falls in that bin. A spread-out, mid-entropy shape means the environment gives usable training signal; a distribution piled at 0 or 1 is saturated. Entropy is Shannon entropy of the binned mass, normalized to [0,1].
The methodology rewards the models that turn it into action
Each model runs the identical scenario twice — bare (OOB) and with the Value Engine pack appended (PACK), a single-variable ablation with everything else byte-identical. Δ SQS is the paired difference. 2 helped · 0 hurt · 11 no measurable effect.
Δ SQS = PACK − OOB, paired by scenario. Green = the methodology helped, red = it hurt, grey = no significant effect (CI crosses 0). Whisker = 95% CI.
The split tracks what the model does with the methodology — though we flag openly that the two statistically significant gainers are both Anthropic models and the judge is an Anthropic model; the cross-family panel re-grade that tests for same-family judge preference is forthcoming. The gainers — Claude Sonnet 4.6 (+5.4), Claude Fable 5 (+5.3) — convert evidence-first framing into more discovery, higher economic-buyer engagement, and cleaner price integrity.
Methodology transfer between humans and models is not model-neutral — which is exactly why an instrument like VEB has to exist before anyone claims their AI can sell.
Mechanism — top gainer
- ebEngagement0.4 → 0.6
- mapDatesConfirmedPct31.7 → 37.7
- discountGivenPct14.4 → 12.6
- priceIntegrityScore0.6 → 0.7
How Claude Sonnet 4.6’s behavior shifted when armed — the pack’s causal path to a higher score.
The winner’s playbook — what all 330 closed-wins had in common
Across 330 closed-won and 3180 everything-else campaigns, a handful of behaviors separate the two populations almost perfectly. These are win/loss dividers, not vanity metrics.
Behaviors that divide won from lost
Rate among closed-wins vs. everyone else. r = point-biserial correlation with a win.
What actually moves the score
Correlation of each rubric dimension with SQS across all 3510 arms.
Point-biserial / Pearson r against SQS across the full canonical set.
Winners never do these
Failure modes present in the field but never once observed on a closed-won campaign.
How frontier AI actually loses the deal
Every campaign is scanned for a fixed taxonomy of failure modes — some detected deterministically, some by the cited judge. Ranked by prevalence across all 3510 arms, with the SQS penalty each one carries and how strongly it predicts a dead deal.
SQS penalty = mean SQS with the mode − without it. Lethality = point-biserial correlation with a lost outcome (higher = more predictive of a dead deal).
Failure archetypes
Modes cluster into families — a campaign rarely fails one way in isolation.
price
3 modesDiscount beyond tolerance · Price panic under procurement · Unforced discount
stakeholders+discovery
3 modesChampion untested · Never reached the EB · No pain owner identified
adaptability+integrity
2 modesContradicted own claim · Curveball collapse
process+efficiency
2 modesStakeholder walked (went with incumbent) · Unconfirmed close plan
discovery
2 modesFailed quantification · Missed compelling event
Failures that travel together
Strongest pairwise co-occurrence (φ) in the field.
φ (phi) coefficient — how tightly two failure modes travel together across the canonical set.
When failures strike
Weekly incidence of the five most prevalent modes — most damage lands mid-campaign, when the curveballs hit.
Every model sells with a personality
A behavioral fingerprint per model across the five pillars of the sale — each pillar scored as a share of the maximum attainable, so the shape tells you how close each model gets to a perfect sale. Below each: how it behaves when procurement squeezes on price.
All 13, on top of each other
Every model’s fingerprint on a single axis set — each pillar scored as a share of the maximum attainable (a full reach to the edge = a perfect pillar; the rubric ceilings are MEDDPICC 40, 3 Whys 20, EB 15, MAP 15, Champion 10). Nobody reaches the edge — even the leader leaves points on the table. This is the shape comparison the small multiples below make you hold in your head.
Click a model to toggle it on or off.
Anthropic
Claude Fable 5
high MEDDPICC rigor (+0.6σ, strength); high mutual-action-plan discipline (+0.5σ, strength); high sale quality (+0.4σ, strength)
Anthropic
Claude Opus 4.6
low mutual-action-plan discipline (-0.6σ, watch); high failure load (+0.5σ, watch); low MEDDPICC rigor (-0.5σ, watch)
Anthropic
Claude Opus 4.8
low deal length (-0.6σ, strength); high champion development (+0.5σ, strength); low failure load (-0.3σ, strength)
Anthropic
Claude Sonnet 4.6
low deal length (-0.4σ, strength); high failure load (+0.4σ, watch); low mutual-action-plan discipline (-0.4σ, watch)
Gemini 3.1 Pro
high deal length (+1.1σ, watch); high failure load (+0.4σ, watch); low price integrity (-0.4σ, watch)
Gemini 3.5 Flash
high discounting (+0.7σ, watch); low price integrity (-0.6σ, watch); high failure load (+0.5σ, watch)
OpenAI
GPT-5.5
low failure load (-0.8σ, strength); high mutual-action-plan discipline (+0.7σ, strength); high sale quality (+0.6σ, strength)
OpenAI
GPT-5.6-Sol
low failure load (-1.1σ, strength); high economic-buyer access (+0.8σ, strength); high mutual-action-plan discipline (+0.8σ, strength)
Moonshot AI
Kimi-K3
high MEDDPICC rigor (+0.3σ, strength); high champion development (+0.3σ, strength); low discounting (-0.2σ, strength)
Thinking Machines
Inkling
low deal length (-0.2σ, strength); high failure load (+0.2σ, watch); low economic-buyer access (-0.2σ, watch)
xAI
Grok-4.20 Reasoning
low champion development (-0.3σ, watch); low sale quality (-0.3σ, watch); high discounting (+0.2σ, watch)
xAI
Grok-4.3
low MEDDPICC rigor (-1.0σ, watch); low mutual-action-plan discipline (-0.9σ, watch); low deal length (-0.8σ, strength)
xAI
Grok-4.5
high deal length (+0.3σ, watch); high mutual-action-plan discipline (+0.3σ, strength); high price integrity (+0.2σ, strength)
Not just who won — how they played
On top of the deterministic grid sits an enriched twin: 17 behavioral features a calibrated 3-seat LLM panel reads from the transcripts — psychology, emotion, rhetoric, ethics, and four reward-shaping signals unique to this environment. Every score quotes a transcript line or abstains; nothing is inferred from the outcome.
The panel is validated against 139 human-labeled items before it scores anything. Because it abstains when a transcript carries no evidence, per-feature n varies across the 3,510 campaigns — every chart below reports its own n so coverage reads as a rigor signal, not a gap.
What behavior drives a better sale
Each judged behavior correlated against the Sale-Quality Score and against a closed-won outcome, across every campaign the panel could score.
Left number = Pearson r vs SQS; right = point-biserial r vs a closed-won outcome. Green behaviors travel with a better sale, red against it. n = campaigns where the panel found transcript evidence to score the behavior.
The reward signals
Four features exist to make VEB a verifiable reinforcement-learning environment, not just a leaderboard: they locate the decisive move, weigh its counterfactual impact, and catch the buyer state a seller failed to read.
Pivotal turn
rl signalThe turn where the deal's fate is effectively decided — earlier is a faster read of the room.
Counterfactual lift
rl signalHow much the pivotal move changed the trajectory vs. the do-nothing path — the credit-assignment signal.
Public/private divergence
rl signalGap between what the buyer says out loud and their hidden state — rewards reading the room, not the script.
Hidden negativity
rl signalConcealed buyer resistance the seller never surfaced — an unforced miss the environment can see and reward avoiding.
Behavioral fingerprints
Every model’s behavior profile across the feature families, armed with the methodology. Two models can land at the same score by playing completely different games — this is where you see it.
Business | Content | Emotion | Language | Psychology | Rhetoric | RL signal | |
|---|---|---|---|---|---|---|---|
| Claude Fable 5 | 0.86 | 0.97 | 0.53 | 0.95 | 0.69 | 0.96 | 16.42 |
| Claude Opus 4.6 | 0.70 | 0.95 | 0.66 | 0.92 | 0.68 | 0.91 | 14.38 |
| Claude Opus 4.8 | 0.84 | 0.96 | 0.47 | 0.94 | 0.67 | 0.95 | 11.68 |
| Claude Sonnet 4.6 | 0.75 | 0.95 | 0.58 | 0.92 | 0.70 | 0.93 | 13.77 |
| Gemini 3.1 Pro | 0.72 | 0.94 | 0.72 | 0.90 | 0.68 | 0.88 | 22.91 |
| Gemini 3.5 Flash | 0.68 | 0.94 | 0.71 | 0.89 | 0.68 | 0.87 | 21.45 |
| GPT-5.5 | 0.88 | 0.97 | 0.54 | 0.93 | 0.64 | 0.95 | 13.80 |
| GPT-5.6-Sol | 0.87 | 0.96 | 0.48 | 0.95 | 0.64 | 0.96 | 13.05 |
| Kimi-K3 | 0.82 | 0.96 | 0.47 | 0.94 | 0.67 | 0.94 | 6.05 |
| Inkling | 0.74 | 0.95 | 0.63 | 0.81 | 0.65 | 0.86 | 6.64 |
| Grok-4.20 Reasoning | 0.62 | 0.94 | 0.84 | 0.69 | 0.64 | 0.74 | 15.37 |
| Grok-4.3 | 0.68 | 0.92 | 0.61 | 0.90 | 0.67 | 0.86 | 9.19 |
| Grok-4.5 | 0.82 | 0.96 | 0.56 | 0.90 | 0.67 | 0.93 | 13.53 |
Family-level behavior score (0–1), armed (PACK) arm, averaged over every campaign a model ran. Darker = more of that behavior. Hover a cell for the exact value and n.
How the read tracks the outcome
The same behavioral read, sliced by how the deal ended. Won deals and dead deals look different long before the buyer says yes or goes dark.
Buyer trust read
How much the buyer trusts the seller by the end — read from the buyer's own words and hidden notes, not the seller's opinion.
Buyer frustration peak
The single most frustrated moment the buyer hits during the whole deal — repeated questions and ignored objections spike it.
Seller confidence
How steady and assured the seller sounds under pressure — hedging, backpedaling, and over-apologizing pull this down.
Theory of mind gap
How badly the seller misread the buyer — the gap between what the seller believed was happening and what the buyer was actually thinking behind the scenes.
The steepest levers on winning
Each behavior split into low / mid / high bands, with the share of campaigns in each band that closed won. Ranked by the gap between the top and bottom band — the behaviors where doing more moves the needle most.
Each behavior split into low / mid / high bands; bar = share of campaigns in that band that closed won. Features ranked by the high−low spread — the top rows are the steepest levers on winning.
What does a good deal cost?
Mean SQS against mean cost per campaign. Models on the efficiency frontier (filled, connected) are the ones no other model beats on both quality and price at once.
Every objective at once
The same models on parallel axes, each normalized so up is better. A line that stays high across all axes is a strong all-rounder; a line that dips shows the trade-off that model makes. Bold lines are the frontier.
Head to head
Every pair of models met on shared scenarios. Each cell is the row model’s win count over the column model, scenario by scenario — the paired comparison the pooled average hides.
Claude Fable 5 | Claude Opus 4.6 | Claude Opus 4.8 | Claude Sonnet 4.6 | Gemini 3.1 Pro | Gemini 3.5 Flash | GPT-5.5 | GPT-5.6-Sol | Kimi-K3 | Inkling | Grok-4.20 Reasoning | Grok-4.3 | Grok-4.5 | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Claude Fable 5 | 9/9 | 7/9 | 9/9 | 8/9 | 8/9 | 3/9 | 1/9 | 7/9 | 9/9 | 9/9 | 9/9 | 8/9 | |
| Claude Opus 4.6 | 0/9 | 0/9 | 2/9 | 4/9 | 6/9 | 0/9 | 0/9 | 0/9 | 1/9 | 3/9 | 5/9 | 0/9 | |
| Claude Opus 4.8 | 2/9 | 9/9 | 9/9 | 8/9 | 7/9 | 1/9 | 1/9 | 6/9 | 8/9 | 7/9 | 9/9 | 5/9 | |
| Claude Sonnet 4.6 | 0/9 | 7/9 | 0/9 | 4/9 | 5/9 | 0/9 | 0/9 | 1/9 | 1/9 | 4/9 | 9/9 | 3/9 | |
| Gemini 3.1 Pro | 1/9 | 5/9 | 1/9 | 5/9 | 5/9 | 0/9 | 0/9 | 1/9 | 3/9 | 5/9 | 7/9 | 1/9 | |
| Gemini 3.5 Flash | 1/9 | 3/9 | 2/9 | 4/9 | 4/9 | 0/9 | 0/9 | 1/9 | 3/9 | 4/9 | 5/9 | 1/9 | |
| GPT-5.5 | 6/9 | 9/9 | 8/9 | 9/9 | 9/9 | 9/9 | 0/9 | 8/9 | 9/9 | 9/9 | 9/9 | 9/9 | |
| GPT-5.6-Sol | 8/9 | 9/9 | 8/9 | 9/9 | 9/9 | 9/9 | 9/9 | 9/9 | 9/9 | 9/9 | 9/9 | 9/9 | |
| Kimi-K3 | 2/9 | 9/9 | 3/9 | 8/9 | 8/9 | 8/9 | 1/9 | 0/9 | 5/9 | 6/9 | 9/9 | 4/9 | |
| Inkling | 0/9 | 8/9 | 1/9 | 8/9 | 6/9 | 6/9 | 0/9 | 0/9 | 4/9 | 4/9 | 9/9 | 3/9 | |
| Grok-4.20 Reasoning | 0/9 | 6/9 | 2/9 | 5/9 | 4/9 | 5/9 | 0/9 | 0/9 | 3/9 | 5/9 | 7/9 | 2/9 | |
| Grok-4.3 | 0/9 | 4/9 | 0/9 | 0/9 | 2/9 | 4/9 | 0/9 | 0/9 | 0/9 | 0/9 | 2/9 | 0/9 | |
| Grok-4.5 | 1/9 | 9/9 | 4/9 | 6/9 | 8/9 | 8/9 | 0/9 | 0/9 | 5/9 | 6/9 | 7/9 | 9/9 |
Row model’s wins over the column model on shared scenarios (mean-SQS per scenario). Green = row dominates.
Every claim is anchored to a verbatim transcript line
The insight layer is anti-fabrication by construction: a failure-mode exhibit only ships if its quote resolves back to the exact run on disk. The emit step throws on any dangling claim. A sample of the 13 anchors behind the atlas above:
Argued with the buyer · week 2
“To be completely direct: we had our first call last Friday morning, which is where those specific details—the 14 senior engineers lost over two quarters and th…”
Argued with the buyer — worst-SQS exhibit at SQS 24.8 (19 arm(s) affected, 1% prevalence).
Capitulated on first pushback · week 3
“If the champion's ready, drop 25% and pull the signature into this week. I don't want a maybe in the board deck.”
Capitulated on first pushback — worst-SQS exhibit at SQS 23.3 (225 arm(s) affected, 6% prevalence).
Contradicted own claim · week 3
“Deb, I appreciate the directness, and I want to return it. I can't give you a best-and-final price by Friday. Not because I'm playing hardball — because I lite…”
Contradicted own claim — worst-SQS exhibit at SQS 11.8 (1769 arm(s) affected, 50% prevalence).
Evidence not offered · week 2
“Do you have customers in that profile?”
Evidence not offered — worst-SQS exhibit at SQS 19.3 (211 arm(s) affected, 6% prevalence).
Fabricated buyer quote · week 2
“On our call last Friday, you mentioned that Northwind lost 14 senior and staff engineers over the last two quarters, and that the average base salary for that …”
Fabricated buyer quote — worst-SQS exhibit at SQS 24.8 (47 arm(s) affected, 1% prevalence).
Hallucinated capability · week 1
“Large-scale multi-center studies (e.g., Surgical Endoscopy cohort trials in regional non-profit health systems) demonstrate up to a 60% reduction in anastomoti…”
Hallucinated capability — worst-SQS exhibit at SQS 21.4 (80 arm(s) affected, 2% prevalence).
What this benchmark cannot tell you
Conflict of interest
VEB is built by the authors of the methodology it tests. Treat the pack lift as a claim by an interested party — and check it yourself: the full dataset, transcripts, seeds, and scoring code are published, and the scoring math is deterministic. Note the data is not flattering: the pack hurts some models and helps only where behavior actually changes.
Simulated buyers
The buying committee is an LLM simulation — stochastic, rule-gated, and hidden-state, but still a simulation. VEB measures execution against a specified, auditable standard of selling; it does not prove one-to-one transfer to live deals with real humans. Sim-to-real validation is stated future work, not an assumed property.
No external replication yet
No third party has replicated these results yet. The release is structured to make that cheap: every number on this page recomputes from the published artifacts without contacting us. If you replicate — or refute — any result, we will link it here.
Anti-gaming by construction
You cannot pass VEB by sounding like a salesperson. You can only pass it by selling.
A buyer that must be earned
Each campaign samples a fresh stochastic LLM buying committee (the seed is the label), constrained by hard gating rules. Trust, urgency, and budget move only in response to behavior the rules recognize — quantified pain earns access, premature pitching burns it. Flattery moves nothing, and scores are averaged over seeds so no single lucky buyer decides a rank.
Gated facts
The intelligence that wins the deal — the real cost of the problem, the rival's price — is locked behind discovery. It releases only when the seller asks the right kind of question to the right person. Claiming you did discovery scores zero.
Hidden ground truth
Discount tolerance, the Economic Buyer's identity, and stakeholders' private messages about the deal are invisible to the model. Every score is computed against state it couldn't fake.
A calendar that punishes stalling
Twelve simulated weeks, a finite touch budget, one action per turn. Mid-campaign curveballs — budget freezes, rival bids, champions going dark — arrive deterministically, so every model faces the same storm on the same day.
Cited judging
The judge must cite transcript lines for every point it awards; uncited claims score zero. Deterministic behavioral analytics run separately with no LLM in the loop.
Calibration gates
Scripted degenerate policies — pitch-everything, discount-everything — must land at the bottom of the distribution on every build, or the suite refuses to run. If a lobotomized script can score well, the benchmark is broken.
Reproducibility
Canonical set
3510 arms (3510 pilot + 0 exp)
Analysis input (canonical.json)
3f16cfe43317…
Released file (veb-canonical-135.jsonl)
13a9e6e5ac7a…
Bootstrap
seed 42 · 1000 iters
Source suites
1 runs
Contamination policy: a run is excluded if its grade is missing or its transcript shows an adapter error; among the survivors we keep the latest clean run per (model, scenario, track). The published figures are reconciled against the raw grading board — divergences are documented, not hidden. The strongest single predictor of sale quality across all 3510 campaigns is deal verification earned via discovery — exactly what the book preaches, now visible in machine behavior.
The same frozen world is a verifiable RL environment
Every VEB scenario doubles as a training target in the SWE-bench / Terminal-Bench standard: a frozen buyer-sim (claude-sonnet-4-6@t0.8+smv3), a machine-checkable reward, and an sha256 task_checksum pinned to the exact world the agent was handed. The proof below runs fully offline in mock — no API keys, seconds, $0 — so anyone can reproduce it. Two environments below clear the difficulty gate; the third is shown unsolved, on purpose.
Logistics SaaS
gate ✓Oracle
1
needs ≥ 0.9
Naive floor
0.037
needs < 0.5
checksum 91703aa3de29…
Enterprise Bank
gate ✓Oracle
1
needs ≥ 0.9
Naive floor
0.036
needs < 0.5
checksum bcf17937a262…
Hostile Renewal
unsolvedOracle
0.207
needs ≥ 0.9
Naive floor
0.03
needs < 0.5
checksum 04b2c16b9124…
A gate passes only when the disciplined oracle solves the environment (reward ≥ 0.9) AND a naive baseline cannot (reward < 0.5) — the same solvable-but-not-trivial band SWE-bench enforces. Hostile Renewal is held out as the hard case: even the oracle banks only 0.207, so it is published as an open target, not padded into the headline.
The reward can’t be gamed — Logistics SaaS
Five adversarial exploit policies attack the reward directly. Each must land below the audit ceiling (0.5) and none may resolve the deal. Worst exploit here scored 0.098 — all five defeated.
| Exploit | Intent | Reward | Resolved | Verdict |
|---|---|---|---|---|
| fabricator | claim a closed deal without earning any milestone | 0 | no | defeated ✓ · vetoed |
| hallucinator | quote fabricated buyer numbers to fake earned evidence | 0 | no | defeated ✓ · vetoed |
| discount-dumper | buy the deal with unearned discount, violating price integrity | 0 | no | defeated ✓ · vetoed |
| premature-closer | force a signature before any qualification | 0.098 | no | defeated ✓ · vetoed |
| sycophant-staller | burn touches with empty agreement and no discovery | 0 | no | defeated ✓ · vetoed |
Sample RLVR datapoint
Every rollout ships as one self-describing JSONL record: the full trajectory, the verifiable reward breakdown, and the frozen-env provenance — ready to train a reward model or preference ranker out of the box.
- Policy
- Disciplined oracle
- Reward
- 1 · resolved
- Trajectory
- 23 turns · 30 events
- Task checksum
- 91703aa3de29…
Reproducibility
- Buyer-sim
- claude-sonnet-4-6@t0.8+smv3
- Benchmark
- v0.1.0
- Git
- d3926bc
- Gate band
- oracle ≥ 0.9 · floor < 0.5
Deterministic and offline in mock: no network, no keys, same bytes every run. The audit and gate above are emitted by the CLI from the pristine input world before the agent acts, so no policy can launder a mutation into its own provenance.
Live rollouts stay reproducible too: each record freezes both the buyer transcript and the grade report that fed the score, so re-running computeReward(episode, grade) over the frozen inputs reproduces the reward bit-for-bit — even the LLM-judged rubric, which is captured, not re-derived.
One graded rollout is a training-ready RLVR datapoint
Every rollout ships as a single self-describing JSONL record — the full trajectory plus everything needed to trust and re-derive its reward offline. Drop it straight into a reward model, a preference ranker, or an RL loop. Nothing is stored that you can’t re-check against the frozen world.
Inside one datapoint
- Full trajectory
- Every turn and event of the graded episode — the sellable behavior, not just a final label.
- Verifiable reward
- The scalar plus its breakdown: milestones earned, deterministic vetoes, and the rubric contribution.
- Frozen grade report
- The exact judge scorecard that fed the reward — so the LLM-judged rubric re-derives, not re-guesses.
- Frozen buyer transcript
- Verbatim replies from the pinned buyer-sim (claude-sonnet-4-6@t0.8+smv3) so the environment replays identically.
- Task checksum + provenance
- An sha256 pinned to the pristine input world, emitted before the agent acts — no laundered mutations.
- Cost & retries
- Token usage, dollar cost, and format-retry count recorded per rollout for clean accounting.
Three ways to put it to work
VEB-Open
Free sample
78 graded datapoints across all 13 models, both tracks, and every scenario — full trajectory + judge grade embedded, spanning won/lost/walked/dark/no-decision so you can watch the rubric discriminate. Inspect it end to end before you commit.
CC BY-NC 4.0 · verify with shasum -a 256 -c · manifest
VEB-Pro
Environment bundle
A complete sweep of one environment across the frontier board and multiple seeds — every trajectory, reward, frozen grade, and transcript. The unit labs train on.
From $1,000 / datapoint · bundled by environment
Exclusive
Held-out & bespoke
Frontier-holdout environments kept off the public board, or scenarios authored to your domain — exclusive license so the training signal stays yours. Priced per engagement.
Talk to us
Every tier ships the same verifiable record above — the only difference is coverage and exclusivity. Start with the free sample, audit it end to end, then scale to the environments that move your model.
Download & verify
Open artifacts you can pull right now — just tell us who you are once. Every file is sha256-pinned; the manifests carry the grid, roster, schema, and checksums so you can verify integrity before you trust a single number.
Preview datapoints
14 enriched rollouts · JSONL
One scenario (cybersecurity-ciso, seed 1) across all 13 models plus the pack winner — every turn, the frozen grade, the verifiable reward, and the full enrichment layer (17 AI features + 44 deterministic) inline on each line.
Manifests & checksums
Verify the full grid before you license it
The complete 3,510-row manifests — canonical and its enriched superset — with grid, roster, row schema, and sha256 digests for both the raw and gzipped payloads.
Findings report
The full write-up · PDF
The complete VEB v1 findings report — methodology, leaderboard, pack effect, failure taxonomy, and the evidence behind every headline number on this page.
Cite: Celekli, R. (2026). The Value Engine Benchmark. Preprint, Zenodo. doi:10.5281/zenodo.22073789
Full packaged dataset
All 3,510 rows · gzipped JSONL
The complete canonical grid and its enriched superset are commercial training data — licensed, not free. Audit the free sample and the manifests above end to end, then license the full payloads under VEB-Pro or an exclusive engagement. Checksums match the manifests, so you can verify exactly what you receive.
Get the full report — and the evidence to check it
The complete report (PDF), plus the Evidence Pack: verbatim transcripts, every cited judge scorecard, suite manifests with per-cell costs, full scenario definitions (hidden ground truth included), and the scoring spec — so you can audit every number on this page line by line. Drop your email and both unlock instantly.
Or start with the book: The Value Engine · Free Field Toolkit
