benchmarks · adventurebench

Can your model play a text adventure?

Adventure Bench measures grounded action interpretation: map a player utterance onto the visible scene’s small action vocabulary, or refuse when the request is not grounded.

the cost curve

The score–cost frontier

Recorded full-run cost · 12 models

Every labeled frontier point is a model for which no cheaper tested model scored as well. The other points are dominated on this release; exact values for every point appear in the table below.

Adventure Bench score versus recorded full-run cost Pareto frontier: Granite Micro → Ministral 3B → Ministral 8B → Ministral 14B → GLM 5.2. 66.2% 73.0% 79.7% 86.4% 93.2% $0.000 $0.032 $0.064 $0.096 $0.128 RECORDED FULL-RUN COST (USD)ADVENTURE BENCH SCORE Granite Micro (ibm-granite/granite-4.0-h-micro): 73.4% at $0.008Ministral 3B (mistralai/ministral-3b-2512): 84.3% at $0.014google/gemma-3-4b-it: 72.5% at $0.020ibm-granite/granite-4.1-8b: 75.8% at $0.020meta-llama/llama-3.1-8b-instruct: 73.6% at $0.020Ministral 8B (mistralai/ministral-8b-2512): 84.4% at $0.022meta-llama/llama-3.2-3b-instruct: 68.9% at $0.023ibm-granite/granite-4.2-8b: 79.5% at $0.023nvidia/nemotron-3.5-lightning: 84.4% at $0.033qwen/qwen3.5-9b: 82.7% at $0.039Ministral 14B (mistralai/ministral-14b-2512): 86.5% at $0.043GLM 5.2 (z-ai/glm-5.2): 90.6% at $0.119 Granite MicroMinistral 3BMinistral 8BMinistral 14BGLM 5.2
Pareto frontierDominated in this release

The frontier is a cost-efficiency view, not a ranking or a statistical claim. Costs are summed from provider-recorded completion evidence; score uncertainty and paired comparisons remain in the table below.

Audited Adventure Bench release 20260903-plicara-v1-expanded. Models are alphabetical by their canonical slug; comparison counts show only paired bootstrap intervals that exclude zero.
ModelScore (95% CI)PassedRecorded costProvider / runtimeParse / transportSeparates
google/gemma-3-4b-it72.5% ; 95% cluster-bootstrap interval (66.8% to 77.9%)531/732$0.020DeepInfra / openrouter:chat-completions-v10 / 0+0 / −7
ibm-granite/granite-4.0-h-micro73.4% ; 95% cluster-bootstrap interval (67.6% to 78.7%)537/732$0.008Cloudflare / openrouter:chat-completions-v13 / 0+0 / −7
ibm-granite/granite-4.1-8b75.8% ; 95% cluster-bootstrap interval (70.1% to 81.1%)555/732$0.020CoreWeave / openrouter:chat-completions-v10 / 0+1 / −6
ibm-granite/granite-4.2-8b79.5% ; 95% cluster-bootstrap interval (74.2% to 84.4%)582/732$0.023CoreWeave / openrouter:chat-completions-v19 / 0+4 / −3
meta-llama/llama-3.1-8b-instruct73.6% ; 95% cluster-bootstrap interval (68.0% to 79.0%)539/732$0.020Groq / openrouter:chat-completions-v10 / 0+0 / −7
meta-llama/llama-3.2-3b-instruct68.9% ; 95% cluster-bootstrap interval (62.7% to 74.6%)504/732$0.023Parasail / openrouter:chat-completions-v172 / 0+0 / −8
mistralai/ministral-14b-251286.5% ; 95% cluster-bootstrap interval (82.1% to 90.6%)633/732$0.043Mistral / openrouter:chat-completions-v10 / 0+6 / −1
mistralai/ministral-3b-251284.3% ; 95% cluster-bootstrap interval (79.8% to 88.5%)617/732$0.014Mistral / openrouter:chat-completions-v10 / 0+5 / −1
mistralai/ministral-8b-251284.4% ; 95% cluster-bootstrap interval (79.8% to 88.8%)618/732$0.022Mistral / openrouter:chat-completions-v10 / 0+6 / −1
nvidia/nemotron-3.5-lightning84.4% ; 95% cluster-bootstrap interval (79.6% to 88.8%)618/732$0.033DeepInfra / openrouter:chat-completions-v12 / 0+5 / −1
qwen/qwen3.5-9b82.7% ; 95% cluster-bootstrap interval (77.7% to 87.3%)605/732$0.039DeepInfra / openrouter:chat-completions-v10 / 0+5 / −1
z-ai/glm-5.290.6% ; 95% cluster-bootstrap interval (87.0% to 93.9%)663/732$0.119DeepInfra / openrouter:chat-completions-v10 / 0+11 / −0

These results are not a ranking. The separates column is +ahead / −behind and counts only paired 95% case-cluster bootstrap intervals that exclude zero. Every comparison whose interval includes zero is unresolved; unresolved comparisons are unresolved.

Release 20260903-plicara-v1-expanded covers 244 cases × 3 repetitions (732 case-repetitions; repetition numbers 1, 2, 3), collected from 2026-09-02 through 2026-09-03 (UTC). Benchmark version: 1.0.0. Open the audited release artifact.

Failure policy: any transport failure invalidates a collection run and prevents it from entering this release. Parsing failures are counted separately after the frozen malformed-output retry and remain visible in the table. See the full methodology.

Every number above regenerates offline from committed raw responses. The release pins every manifest and response hash plus dataset f38e4650ecbe… and prompt a5b12544f699…. Audit it with make site-check RELEASE_ID=20260903-plicara-v1-expanded.

Per-tag passed/total coverage for the audited release. Columns are models in alphabetical canonical-slug order.
Tag google-gemma-3-4b-it ibm-granite-granite-4-0-h-micro ibm-granite-granite-4-1-8b ibm-granite-granite-4-2-8b-20260831 meta-llama-llama-3-1-8b-instruct meta-llama-llama-3-2-3b-instruct mistralai-ministral-14b-2512 mistralai-ministral-3b-2512 mistralai-ministral-8b-2512 nvidia-nemotron-3-5-lightning qwen-qwen3-5-9b z-ai-glm-5-2
abbreviation12/3012/3030/3021/3024/3012/3027/3013/3024/3021/3021/3027/30
absent-object18/3012/3027/3021/3012/3012/3030/3030/3030/3030/3030/3030/30
adverb21/2121/2121/2121/2121/2118/2121/2121/2121/2121/2121/2121/21
ambiguity9/3012/3012/3021/3018/3015/3030/3018/3016/3024/3027/3030/30
classic-command30/4827/4821/4833/4829/4818/4844/4844/4839/4840/4839/4845/48
compound24/2727/2718/2727/2727/2721/2724/2722/2724/2721/2720/2726/27
direction-as-place42/4542/4542/4542/4542/4542/4545/4542/4542/4542/4545/4545/45
exact-verb36/3636/3636/3636/3636/3636/3636/3636/3636/3636/3636/3636/36
flavor-noun6/429/423/426/420/423/4217/4218/4211/4218/4215/4214/42
frustration9/1212/1212/1212/129/129/1212/1212/1212/129/1212/1210/12
full-sentence24/2424/2421/2424/2424/2424/2421/2424/2424/2424/2421/2424/24
go-to-place9/129/129/129/129/129/1211/129/129/127/1212/129/12
greeting3/99/99/90/90/90/99/99/99/99/99/98/9
guess-the-noun6/1212/1212/126/1212/1212/1212/1212/1212/129/1212/1212/12
help-seeking9/1512/159/1515/159/159/1515/1515/1515/159/1512/1515/15
meta15/3018/3027/3030/3027/309/3027/3027/3030/3027/3030/3027/30
missing-preposition18/2727/2727/2727/2727/2727/2724/2724/2727/2727/2727/2727/27
multi-object15/1515/153/1515/1515/1515/1515/1513/1515/1515/1512/1515/15
nonsense9/129/129/1212/1212/126/1212/1210/129/1212/1212/1212/12
out-of-vocab60/12642/12645/12663/12650/12648/12695/12698/12680/12697/12671/12697/126
paraphrase42/5151/5148/5148/5148/5142/5145/5144/5145/5145/5148/5151/51
politeness18/1818/1818/1818/1815/1818/1815/1818/1818/1815/1818/1816/18
pronoun12/2715/2715/2721/2718/2715/2721/2721/2719/2724/2724/2725/27
question42/4536/4545/4545/4536/4530/4535/4545/4545/4542/4542/4545/45
relative-direction3/123/123/126/123/123/129/123/126/123/123/1212/12
story-mode9/90/99/99/99/96/99/99/99/99/99/99/9
synonym60/6363/6363/6363/6363/6360/6360/6360/6363/6357/6360/6360/63
typo36/3636/3633/3636/3633/3636/3630/3631/3636/3634/3633/3636/36
verb-alias15/1515/1515/1515/1515/1515/1515/1515/1515/1515/1515/1515/15

Per-tag cells are numerators and denominators, not pooled rankings; they show which grounded mapping and calibration patterns each audited run passed.

Repository Raw evidence Dataset Card Re-run it ← All benchmarks